Uncategorized Data science Screen Test by mfh.officials@gmail.com Jan 7, 2025 0 Comment Data science Data science Screening test 1 / 20 Which of the following is TRUE about ensemble methods? They combine multiple weak models to create a stronger model They are less prone to overfitting than single models They are always more computationally expensive than single models They cannot be used with decision trees 2 / 20 Which of the following metrics is used to evaluate clustering models? Precision ROC-AUC Silhouette score F1-score 3 / 20 The "curse of dimensionality" refers to: The difficulty in finding a suitable machine learning model The difficulty of visualizing data in high-dimensional spaces The tendency of high-dimensional data to become sparse The increasing complexity of models with more features 4 / 20 In random forests, what is the primary advantage over a single decision tree? It uses more training data It always uses shallow trees It reduces variance and improves accuracy by averaging predictions It is more interpretable 5 / 20 Which of the following techniques is used for dimensionality reduction? Principal Component Analysis (PCA) K-means clustering Naive Bayes Decision Trees 6 / 20 Which of the following is the most appropriate way to deal with missing values in a dataset? Replace missing values with the mean or median of the column Replace missing values with zeros Drop all rows with missing values Use a model-based imputation method 7 / 20 Which of the following is an example of unsupervised learning? Logistic Regressi Linear Regression Decision Trees Principal Component Analysis (PCA) 8 / 20 In a time series forecasting problem, which of the following is most commonly used to check for stationarity? Shapiro-Wilk test Augmented Dickey-Fuller (ADF) test ACF/PACF plots Durbin-Watson test 9 / 20 Which evaluation metric is most appropriate for imbalanced classification problems? Precision and Recall Accuracy Mean Squared Error R-squared 10 / 20 Which of the following is a disadvantage of the k-nearest neighbors (KNN) algorithm? Answer: B) It requires a large amount of training data It requires a large amount of training data It is difficult to interpret It cannot handle multi-class classification It assumes linearity in the data 11 / 20 Which of the following is a hyperparameter for the k-means clustering algorithm? Number of clusters (k) Activation function Learning rate Regularization strength 12 / 20 Which of the following libraries is primarily used for deep learning? Matplotlib scikit-learn TensorFlow pandas 13 / 20 What is the purpose of the Adam optimizer in neural networks? To perform backpropagation To calculate the gradient of the loss function To adjust the learning rate during training To reduce the loss function 14 / 20 Which of the following algorithms is most appropriate for predicting a continuous outcome variable? K-means clustering Linear regression Decision tree classification K-nearest neighbors (classification) 15 / 20 Which of the following is a key assumption of the linear regression model? Multicollinearity Non-linearity between independent and dependent variables Homoscedasticity Independence of dependent variables 16 / 20 2. In a decision tree, the split criterion is typically based on: Variance Root mean squared error Information gain or Gini impurity Sum of squared errors 17 / 20 In the context of model evaluation, what does the "ROC curve" stand for? Residual Output Curve Recurrent Operations Curve Receiver Operating Characteristic Curve Root Output Curve 18 / 20 Cross-validation is primarily used to: Reduce overfitting and assess model performance Split data into training and testing sets Tune hyperparameters Visualize data 19 / 20 What is the "bias-variance tradeoff"? Increasing model complexity does not affect bias or variance Increasing model complexity reduces variance and increases bias Increasing model complexity reduces both bias and variance Increasing model complexity reduces bias and increases variance 20 / 20 What is the purpose of regularization in machine learning? To reduce overfitting by penalizing large coefficients To speed up training To increase model complexity To improve model interpretability Your score is
Leave a Reply Cancel replyYour email address will not be published. Required fields are marked *Comment Name* Email* Save my name, email, and website in this browser for the next time I comment.
Leave a Reply