Uncategorized Data science Screen Test by mfh.officials@gmail.com Jan 7, 2025 0 Comment Data science Data science Screening test 1 / 20 Which of the following algorithms is most appropriate for predicting a continuous outcome variable? K-means clustering K-nearest neighbors (classification) Decision tree classification Linear regression 2 / 20 Which of the following is TRUE about ensemble methods? They are less prone to overfitting than single models They cannot be used with decision trees They are always more computationally expensive than single models They combine multiple weak models to create a stronger model 3 / 20 The "curse of dimensionality" refers to: The difficulty in finding a suitable machine learning model The increasing complexity of models with more features The tendency of high-dimensional data to become sparse The difficulty of visualizing data in high-dimensional spaces 4 / 20 Which of the following is a hyperparameter for the k-means clustering algorithm? Activation function Regularization strength Learning rate Number of clusters (k) 5 / 20 Which of the following techniques is used for dimensionality reduction? Naive Bayes Principal Component Analysis (PCA) K-means clustering Decision Trees 6 / 20 Cross-validation is primarily used to: Visualize data Tune hyperparameters Split data into training and testing sets Reduce overfitting and assess model performance 7 / 20 In random forests, what is the primary advantage over a single decision tree? It is more interpretable It always uses shallow trees It uses more training data It reduces variance and improves accuracy by averaging predictions 8 / 20 What is the purpose of regularization in machine learning? To speed up training To improve model interpretability To reduce overfitting by penalizing large coefficients To increase model complexity 9 / 20 What is the purpose of the Adam optimizer in neural networks? To reduce the loss function To calculate the gradient of the loss function To adjust the learning rate during training To perform backpropagation 10 / 20 2. In a decision tree, the split criterion is typically based on: Information gain or Gini impurity Sum of squared errors Variance Root mean squared error 11 / 20 In a time series forecasting problem, which of the following is most commonly used to check for stationarity? Shapiro-Wilk test Durbin-Watson test Augmented Dickey-Fuller (ADF) test ACF/PACF plots 12 / 20 Which of the following is a disadvantage of the k-nearest neighbors (KNN) algorithm? Answer: B) It requires a large amount of training data It cannot handle multi-class classification It requires a large amount of training data It is difficult to interpret It assumes linearity in the data 13 / 20 In the context of model evaluation, what does the "ROC curve" stand for? Residual Output Curve Root Output Curve Receiver Operating Characteristic Curve Recurrent Operations Curve 14 / 20 Which of the following is an example of unsupervised learning? Logistic Regressi Principal Component Analysis (PCA) Linear Regression Decision Trees 15 / 20 Which of the following is the most appropriate way to deal with missing values in a dataset? Replace missing values with the mean or median of the column Replace missing values with zeros Drop all rows with missing values Use a model-based imputation method 16 / 20 Which evaluation metric is most appropriate for imbalanced classification problems? Precision and Recall Mean Squared Error R-squared Accuracy 17 / 20 Which of the following is a key assumption of the linear regression model? Multicollinearity Homoscedasticity Non-linearity between independent and dependent variables Independence of dependent variables 18 / 20 Which of the following metrics is used to evaluate clustering models? F1-score Silhouette score Precision ROC-AUC 19 / 20 What is the "bias-variance tradeoff"? Increasing model complexity does not affect bias or variance Increasing model complexity reduces bias and increases variance Increasing model complexity reduces both bias and variance Increasing model complexity reduces variance and increases bias 20 / 20 Which of the following libraries is primarily used for deep learning? Matplotlib pandas scikit-learn TensorFlow Your score is
Leave a Reply Cancel replyYour email address will not be published. Required fields are marked *Comment Name* Email* Save my name, email, and website in this browser for the next time I comment.
Leave a Reply