Browse all practice questions for the Casualty Actuarial Society MAS-1 Practice Exam. Search by topic, open any question and review its full explanation, then test yourself in the practice quiz.

Casualty Actuarial Society MAS-1 Practice Exam course image
More practice questions

These questions are part of the practice quiz. Start practicing

  • True or false: Local regression is a memory-based procedure.
  • The MVUE of the scale parameter beta in Gamma(shape=alpha, scale=beta) is which expression?
  • Both ridge and lasso regression are regularized methods.
  • Best subset selection requires fitting all (2 choose p) models for each possible combination of p predictors.
  • If X_(k) is the kth order statistic from an iid sample from Uniform(0, θ), which distribution does X_(k) follow?
  • In smoothing splines, setting lambda to zero corresponds to no penalty for roughness.
  • Which inequality defines the likelihood ratio test critical region?
  • What is the likelihood ratio critical region for testing H0 against H1?
  • For an unbiased estimator, the MSE is always equal to the variance.
  • Regarding a simple linear relationship, if the irreducible error is zero (e = 0), the 95% confidence interval is equal to the 95% prediction interval.
  • Dimension reduction involves projecting into a lower-dimensional subspace using M linear combinations.
  • With E(N)=150, Var(N)=100, z=1.96, what is the ILT lower bound for the number of deaths?
  • Increasing the significance level would decrease the power of a test.
  • In kernel density estimation using a Gaussian kernel, the width of the neighborhood is infinite.
  • In k-fold cross-validation, the model is fitted a total of k times.
  • What is Mallows Cp equation?
  • Given the minimal path sets {1,2,5}, {1,3,4}, {2,3,5}, {3,4,5}, which of the following is a minimal cut set?
  • Conditional on θ, X follows a Normal distribution with mean θ and variance 100^2.
  • True or false: a small deviance indicates a poor fit.
  • What happens to training mean squared error as model flexibility increases?
  • How can you identify the mode using a probability density function?
  • Which methods guarantee a nested sequence of models as predictors are added or removed?
  • In smoothing spline models, increasing lambda affects the bias-variance tradeoff. Which statement is correct?
  • Which statement about bootstrapping and cross-validation is false?
  • A distribution with support depending on theta cannot be a member of the standard exponential family.
  • When should you use pooled variances for confidence intervals (or tests)?
  • The first PC is the line in p-dimensional space that is closest to the observations.
  • SSE equal to 0 indicates overfitting.
  • What makes a good argument for choosing LOOCV over 5-fold CV?
  • A 5-state Markov chain with two classes {0,1,3} and {2,4} is ergodic.
  • What is the test statistic for testing the significance of a single parameter in a regression model?
  • What does removing outliers do to a linear regression model?
  • In a one-dimensional symmetric random walk, where the probability of moving in either direction is 0.5, are all states recurrent?
  • Which distribution is associated with the canonical link that is inverse?
  • Using all possible PCs provides the best understanding of the data.
  • N(t) must be greater than or equal to 0 is a property of counting processes.
  • Best subset selection requires fitting all possible subset models, a total of 2^p models.
  • Adjusted R^2 equal to 1 indicates overfitting.
  • The first principal component direction of the data is the axis along which the observations vary the most.
  • What happens to the variance-covariance matrix when implementing the quasi-likelihood method?
  • Which formula correctly expresses Var(X+Y) accounting for dependence?
  • For a distribution in the exponential family, E[Y] equals negative derivative ratio - c'(θ) / b'(θ).
  • Power is defined as 1 minus the probability of a Type II error.
  • In simple linear regression, a random pattern in the scatterplot of y against x indicates that R^2 is near zero.
  • What's the formula for Cov(X,Y)?
  • Poisson regression assumes that the mean equals the variance of the response variable.
  • What is the canonical link function for Poisson regression?
  • LOOCV is a special case of k-fold cross-validation.
  • Using an alternative fitting procedure will likely improve prediction accuracy.
  • Which statement is NOT an assumption when using pooled variances for two-sample tests?
  • When forming a confidence interval for the difference between two means with known variances, which statistic is used?
  • If X ~ Uniform(m, n), what is the distribution of (X | X > pi_q)?
  • In the Tweedie family, p in (1,2) corresponds to which distribution?
  • In general, LOOCV requires fitting a model for a total of n times.
  • PCA serves as a tool for data visualization.
  • In a Markov chain, a state that cannot be left once entered is called an absorbing state.
  • In a Markov chain, a state whose probability of returning is less than 1 is called a
  • In Lasso regression, as the regularization parameter lambda increases, what happens to the number of predictors selected?
  • In simple linear regression, the least squares line passes through the point (x-bar, y-bar).
  • In actuarial notation, the symbol a_double_dot_40 denotes:
  • Increasing model flexibility decreases variance.
  • If the number of claims follows a Poisson distribution with rate lambda, what distribution describes the waiting time until the first claim?
  • The Tweedie distribution is particularly useful when data include zeros and continuous positive values that can be viewed as which distribution?
  • In a linear model with an intercept and p explanatory variables, the leverage for each observation must be between which values?
  • True or false: Ridge regression is less flexible and thus results in an improved prediction accuracy when its decrease in variance is less than its increase in squared bias.
  • Ridge regression outperforms lasso when the response is a function of many predictors, all with coefficients of roughly equal size.
  • GAMs are a useful representation if we are interested in inference, since you can examine the effect of the predictor variables on the response while holding all of the other predictor variables constant.
  • Lasso regression tends to outperform ridge regression in terms of bias, variance, and MSE.
  • To determine asymptotic unbiasedness, which condition must hold?
  • For a negative binomial distribution with parameter r, the MVUE of the scale parameter is:
  • Collinearity reduces the accuracy of the estimates of regression coefficients and may make it harder to reject the hypothesis that beta_j = 0.
  • Which of the following statements best describes a non-homogeneous Poisson process?
  • If two statistics have the same mean, the one with smaller variance is called the efficient estimator.
  • Deviance in generalized linear models and the chi-square distribution: deviance follows a chi-square distribution for all models in the exponential family.
  • Increasing the significance level would increase the probability of a Type I error.
  • It is possible to directly estimate a model's test error using a validation set or cross-validation.
  • The logit model is appropriate when the response variable is binary.
  • In ridge regression, the sum of squares of beta is bounded above by s. Which statement best captures this constraint?
  • Which of the listed modeling procedures performs variable selection?
  • When forming a confidence interval for a proportion, which statistic is used?
  • PCA can be used for data visualization.
  • In a Markov chain, positive recurrence is a property that holds for all states in a given communicating class.
  • The irreducible error variance in this model is 100^2.
  • In PCR, is it recommended to standardize each predictor prior to generating principal components?
  • Which description matches the constrained form used in the lasso alternative objective?
  • Which statement correctly describes the relationship between MSE and the true parameter?
  • To form a confidence interval for the ratio of variances between two populations, which statistic is used?
  • If a simple linear model for a probability predicts values outside the [0,1] interval, what is the standard corrective approach?
  • Using a canonical link function in a GLM, are the estimates unbiased or biased?
  • A probit link is a valid alternative to the logistic link for binary outcomes.
  • SSE equals zero indicates overfitting.
  • In a series system, the system functions only if every component is functioning.
  • Which distribution and link function should be used for slices of pizza sold at a convenience store based on distance to the city center?
  • In the Poisson distribution, which statement is true about the relationship between the mean and the variance?
  • A series system functions only when all components function.
  • If the hazard rate function decreases with x, the distribution has a heavy tail.
  • In PCA, the first few principal components are often sufficient to get a good understanding of the data.
  • Which model form will both have a discontinuity in the fitted curve and most likely overfit the data when predicting height from shoe size?
  • When forming a confidence interval for a population variance, what kind of statistic is used?
  • R^2 equal to 1 indicates overfitting.
  • Which statement about the kth order statistic from Uniform(0, θ) is correct?
  • True or false: A large value of Mallows Cp indicates a model with a high test error.
  • Which statement about lasso regression compared to ordinary least squares is true?
  • A saturated model has a deviance of zero.
  • When forming a confidence interval for a mean with a known variance, which statistic is used in the calculation?
  • What is the formula for a Pearson residual?
  • The deviance is defined as a measure of distance between saturated and fitted model.
  • what is the neyman-pearson theorem for hypothesis testing?
  • The incremental variance explained by adding another principal component decreases as more components are included.
  • In a three-state Markov chain with states 0, 1, 2 and starting in 0, what is the formula for the expected number of steps to return to state 0?
  • Kernel density estimation is used to estimate which component of a distribution?
  • Ridge regression uses an L2 penalty on coefficients.
  • Which statement is true about backward stepwise selection?
  • PCR assumes that the directions in which features show the most variation are the directions that are associated with the target.
  • Regarding kernel density estimation: the larger the bandwidth, the smoother the estimated pdf is.
  • In a fitted values vs residuals graph, heteroscedasticity is indicated by which pattern?
  • k-fold cross-validation has higher variance than LOOCV when k<n.
  • A consistent estimator cannot be biased.
  • Rank the following tools by flexibility in descending order: spline, linear regression, ridge regression.
  • How many minimal cut sets are there for a random graph with n nodes?
  • Evaluating the correlation matrix of predictor variables is a reliable method to detect collinearity.
  • Is KNN an example of supervised or unsupervised learning?
  • In ridge regression, the shrinkage penalty is applied to all coefficient estimates except for the intercept.
  • In the gambler's ruin scenario, what is Ben's expected final wealth when the total is 75 and the win probability is 0.5?
  • What is the test statistic for testing the equality of two variances?
  • Poisson regression models incorporate a logarithmic link function.
  • In least squares regression, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • Which statement is true about the scale behavior of ridge regression?
  • True or false: The model containing all predictors will always have the smallest residual sum of squares and largest R-squared.
  • Mallows Cp and AIC are proportional to each other; in fact, they are equal.
  • What is the canonical link for the inverse Gaussian distribution?
  • Cross-validation is used to measure the accuracy of a parameter estimate.
  • Performing k-fold cross validation requires fitting a model for a total of k times.
  • A Poisson process with a constant rate is called what?
  • If the branching process starts with n individuals, the extinction probability is π0^n.
  • True or false: Local regression should not be used in a high-dimensional setting.
  • A large value of a leverage indicates the presence of an outlier.
  • Natural splines, regression splines, smoothing splines, local regression, polynomial regression, and step functions are all types of models that can be used as building blocks for GAMs.
  • Best subset selection results in a nested set of best models with different numbers of predictors.
  • LOOCV bias is lower than k-fold CV bias.
  • Is PCA an example of supervised or unsupervised learning?
  • If X ~ Exp(λx) and Y ~ Exp(λy) are independent, what is E[min(X,Y)]?
  • The smoothness of a continuous predictor variable in a GAM can be summarized by degrees of freedom.
  • In a residuals vs fitted values plot, heteroscedasticity is indicated by which pattern?
  • In ridge regression, which statement is true?
  • What is the expected value of the top q% of losses?
  • Which statement describes a common effect of high dimensionality on MSE estimates?
  • True or False: The F-statistic used to test a predictor in regression is computed as the ratio of its mean square to the mean square error from the model with the most predictors.
  • If there is a correlation among the error terms, then the estimated standard errors will tend to underestimate the true standard errors.
  • What is the canonical link for the gamma/exponential distribution?
  • What is the canonical link for the Poisson distribution?
  • In Lasso regression, as lambda increases, what happens to the variance of the predictions?
  • True or false: A small value of span s results in a global fit for a local regression.
  • How do you standardize a residual?
  • Which statistic would you use to form a confidence interval for the difference of two means when the variances are unknown?
  • Which of the following statements is true about one-dimensional and two-dimensional symmetric random walks?
  • Which statement about leverages in a linear model with intercept and p explanatory variables is correct?
  • True or false: Mallows Cp and AIC are proportional to each other.
  • Bias refers to the error arising from the method's sensitivity towards the training data set.
  • In nested models, the deviance is useful for testing the significance of explanatory variables.
  • In a hierarchical normal model where X|θ ~ N(θ, 100^2) and θ ~ N(800, 50^2), what is E[X]?
  • In the context of the material, when estimating an exponential distribution, the MLE of the mean equals the sample mean.
  • If theta-hat is unbiased and efficient, then theta-hat is the MVUE.
  • When applying the quasi-likelihood approach to a GLM, does it change the point estimates of the coefficients?
  • It is possible to estimate test error by adjusting training error to account for bias due to overfitting.
  • For a Negative Binomial distribution with parameter r, which expression is the MVUE of the scale parameter?
  • What is the canonical link for the normal distribution?
  • In quasi-likelihood, how is the variance-covariance matrix adjusted?
  • LOOCV requires fitting a model a total of n times.
  • In a Markov chain, a state that is guaranteed to be revisited eventually is called a
  • In a regression with p predictors and an intercept, the trace of the hat matrix (sum of leverages) equals:
  • What is the MVUE of a normal distribution with a defined mean?
  • When given a CDF of a distribution, how do you obtain the CDF of the kth order statistic Y_(k)?
  • What is the matrix S defined as in the method for transient absorption probabilities?
  • Ridge regression cannot set any coefficient exactly to zero.
  • Some regularization methods can also perform variable selection by estimating coefficients to be precisely zero.
  • Which of the following increases monotonically as model flexibility increases?
  • A chain with only one class is called an irreducible chain.
  • Unlike the validation set approach, the k-fold cross-validation approach uses all observations to train the model.
  • In regression context, if there is no linear relationship between x and y, the scatterplot will typically show a random pattern.
  • LOOCV corresponds to k-fold cross-validation with k equal to n.
  • In simple linear regression, R^2 equals the square of the sample correlation coefficient r. Which option is true?
  • X ~ Exp(theta). any loss over 10,000 will result in a claim payment of only 10,000 due to policy limits. you observe 4 claim payments: 1000, 3100, 7500, 10000. how would you calculate L(theta)?
  • PCA finds a low dimension representation of a dataset that contains as much variation as possible.
  • True or false: larger values of lambda result in greater effective degrees of freedom for the model.
  • N(t) doesn't have to be an integer.
  • Mallows Cp is an unbiased estimate of the test MSE if calculated using an unbiased estimate of variance.
  • In a lifetime model where lifetimes are i.i.d. exponential with mean θ, the expected value of the k-th order statistic X_(k) is equal to which expression?
  • In a local regression model, increasing the span s will typically produce what effect on the fitted curve?
  • Under a saturated model, the predicted value for a given observation is
  • The efficiency of an estimator is defined as the Rao-Cramer lower bound divided by the estimator's variance.
  • Using an alternative fitting procedure will likely result in a simpler model.
  • In a parallel system, the system fails only when all components fail.
  • If an explanatory variable is uncorrelated with all other explanatory variables, the corresponding VIF would equal what?
  • Shrinkage reduces variance at the cost of a small increase in bias.
  • If Y is complete sufficient for theta and g(Y) is unbiased for theta, then g(Y) is the MVUE with the smallest variance.
  • Deviance can be used to test the significance of explanatory variables in nested models.
  • Deviance is minimized to obtain the best-fitting generalized linear model; in general, lower deviance indicates a better fit.
  • What best describes the type of problem Angela is solving by clustering shoppers to target ads?
  • In ridge regression, increasing the tuning parameter lambda strengthens the penalty and shrinks coefficients toward zero.
  • True or false: since training error can be a poor estimate of the test error, RSS and R-squared are not suitable for selecting the best model.
  • Main objective of ridge regression?
  • Do the beta_hat values of a ridge regression procedure provide unbiased estimators of the corresponding beta model parameters?
  • If theta-hat is the MVUE, theta-hat is efficient.
  • An unbiased estimator is considered a consistent estimator if the variance of the estimator converges to 0 as n approaches infinity.
  • Which of the following is generally considered unsupervised learning?
  • For a given dataset, the number of variables in a Lasso regression model will always be greater than or equal to the number of variables in a Ridge regression model.
  • The cumulative proportion of variance explained increases as more PCs are added.
  • What is the probability that an observation is not selected for a bootstrap sample?
  • Explanatory variables for Poisson regression besides exposure can be either continuous or categorical.
  • Regularized regression methods include ridge and lasso, and their purpose is to prevent overfitting.
  • Which of the following expresses ridge regression as a constrained optimization problem?
  • True or false: all states in an irreducible Markov chain are recurrent.
  • Residual sum of squares is monotonic with respect to the number of predictors.
  • In the Tweedie family, p = 3 corresponds to which distribution?
  • Variance refers to the error arising from the assumptions made in the statistical learning tool.
  • Is a regression tree an example of supervised or unsupervised learning?
  • Which expression expresses E[Y] for a distribution in the exponential family?
  • The deviance for normal distributions is proportional to the residual sum of squares.
  • What is the typical statement about removing outliers on model fit?
  • Simon uses a statistical learning method to estimate the number of ears of corn produced per acre. He applies the same method to multiple training data sets and results are similar but not identical. What best describes this method?
  • Which statement about Mallows Cp and AIC is supported by the material?
  • In simple linear regression, does the choice of explanatory variable x affect the total sum of squares?
  • For modeling hourly bike-sharing usage by day of week, which distribution and link function are most appropriate?
  • In binomial data, overdispersion manifests as the observed variance exceeding the binomial variance, which is expressed as which formula?
  • What is the symbol used for the expected number of time periods a chain is in state 2 given the chain starts in state 1?
  • In ridge regression, which parameter is not subject to shrinkage?
  • The fewer positive raw moments that exist, the greater the tail weight.
  • In Lasso regression, as lambda increases, the squared bias of the parameters in the model tends to
  • What represents the moment generating function of Y evaluated at t=1, My(1)?
  • Deviance is a useful measure of goodness of fit for all models in the exponential family.
  • Power parameter in the Tweedie family that corresponds to a gamma/exponential distribution.
  • In the given model, E[X] equals E[θ].
  • What is a commonly used model for times to failure (or survival times)?
  • When using the quasi-likelihood approach, is the variance-covariance matrix scaled by an extra dispersion parameter?
  • What does the likelihood ratio test (LRT) test?
  • A parallel system functions as long as one of the components functions.
  • K-fold validation has an advantage over LOOCV in variance reduction.
  • Leave-one-out cross-validation is a special case of k-fold cross-validation where k equals the number of observations.
  • Does the quasi-likelihood method change the coefficient estimates?
  • In kernel density estimation, increasing the bandwidth reduces variance but can increase bias; the statement about smoother pdf holds true when bandwidth is larger.
  • Using the ILT to price life insurance policies, the lower bound for the number of deaths during a period is given by which expression?
  • Is boosting an example of supervised or unsupervised learning?
  • Var(S) for S = sum_{i=1}^N X_i with N ~ Poisson(lambda) and i.i.d. X_i is equal to lambda * E[X^2].
  • Which statement about k-fold cross-validation is true?
  • What is MSE(estimator)?
  • How many of the modeling techniques perform dimension reduction: lasso, PLS, PCA, ridge?
  • Which process is characterized as a counting process with integer-valued counts?
  • In the gambler's ruin scenario with total wealth 75, Ben starts with 40 and Allison with 35. Which expression correctly computes the expected final wealth of Ben?
  • Which norm is used in the penalty term for lasso regression?
  • In least squares LOOCV, the LOOCV error can be computed from a single fitted model using residuals and leverage.
  • Are all Poisson processes characterized by stationary and independent increments?
  • In PCA, the third principal component is orthogonal to the first principal component.
  • In Poisson regression, the Pearson residual is computed as (y - mu_hat) / sqrt(mu_hat). What does mu_hat represent in this context?
  • The training MSE decreases as model flexibility increases.
  • The sum of the leverages across all observations must equal the number of explanatory variables.
  • If all regression errors are identically zero, what is the R-squared value?
  • The deviance is defined as a measure of distance between saturated and fitted model.
  • True or false: as lambda increases from 0 to infinity, the effective degrees of freedom decrease from n to 2.
  • A gamma random variable with alpha = 2 and theta = 1 to be when simulating random variables?
  • When forming a confidence interval for the difference between two means with unknown variances, which statistic is used?
  • Ridge regression objective can be formulated as minimizing SSR subject to which constraint?
  • In the formula f_{2,4} = (s_{2,4} - δ_{2,4}) / s_{4,4} used to compute a probability, what does δ represent?
  • What is the MVUE of sigma^2 for a normal distribution with unknown mean and unknown variance?
  • In a smoothing spline model fit to data, what happens to bias as the tuning parameter lambda increases?
  • Before applying ridge regression, predictors should be standardized because ridge is not scale invariant.
  • In generalized linear models, the statement 'the saturated model has the highest possible deviance' is true or false?
  • After standardizing the predictors, which statement about the PLS first direction is correct?
  • PCR is useful for performing feature selection.
  • In the exponential family, the expression for E[Y] can be written as E[Y] = - c'(θ) / b'(θ).
  • The smallest possible value of a leverage is 0.
  • What is the formula for pseudo R-squared?
  • The Excel function CHISQ.DIST.RT can be used to compute the p-value for a chi-square statistic.
  • In computing the first direction, PLS places the highest weight on the variables that are most strongly related to the response.
  • When a lognormal X ~ Lognormal(mu, sigma^2) is scaled by a positive constant c, which distribution describes cX?
  • In ridge regression, what happens to variance as the budget parameter s increases?
  • To determine if a function should be used as a link function for a GLM, check if the function is monotone and differentiable.
  • True or false: The critical region of a hypothesis test is determined by the significance level and not by the sample observations.
  • True or false: We should choose a model with a low training error when selecting the optimal model.
  • Pr(X>x) with a conditional distribution? (Law of Total Probability)
  • PCA selects low-dimensional linear surfaces to maximize the captured variance.
  • For a natural cubic spline, what is the number of degrees of freedom?
  • R^2 is the ratio of the regression sum of squares to the total sum of squares.
  • Which statement about LOESS span and smoothing is true?
  • Lasso regression is able to perform variable selection by forcing some coefficients to be exactly zero.
  • If both classes were transient, after some time, the chain would not be in either class.
  • How do you identify overdispersion in a model?
  • The cumulative proportion of variance explained cannot decrease when more PCs are added.
  • For paired observations, which statistic is used to form a confidence interval for the difference in means?
  • In a Galton-Watson branching process with offspring probabilities P_j, the extinction probability π0 satisfies which equation?
  • Which of the following best describes the objective of lasso regression?
  • When comparing two means assuming known variances for both populations, which statistic is used to form the confidence interval?
  • In dummy coding for a categorical predictor with four levels, how many dummy variables are needed if one category is used as the baseline?
  • PLS identifies new features in an unsupervised way by approximating the original predictors, similar to PCA.
  • In the Tweedie family, p = 2 corresponds to which distribution?
  • Ridge regression tends to shrink coefficient estimates toward zero and typically does not set any coefficients exactly to zero.
  • Deviance is a measure used to assess the quality of fit for nested models.
  • Which statement best captures the meaning of stationary and independent increments?
  • The standard error of regression uses degrees of freedom equal to n-2.
  • If state 2 is positive recurrent, then state 4 must be positive recurrent.
  • How is Greedy Algorithm A described for optimization problems?
  • When deciding between a regression spline and local regression, which component must be considered for a regression spline but not for local regression?
  • A logit model applies when the response variable counts the number of events occurring.
  • The estimated variance of a coefficient in OLS is given by which expression?
  • what is the excel equation for solving for the pdf in gaussian kernel estimation?
  • Greedy Algorithm A is commonly used for which class of optimization problems?
  • Which sequence leads to the asymptotic variance of θ?
  • LOOCV tends to overestimate the test error rate in comparison to validation set approach.
  • Poisson regression models are capable of handling varying exposure by allowing the exposure term to differ across observations.
  • K-fold validation has a computational advantage over LOOCV when k < n.
  • Which of the following is NOT listed as a potential cause of unreliable mean squared error estimates?
  • What is the formula for the degrees of freedom in a chi-square goodness-of-fit test with g groups and p estimated parameters?
  • What is the canonical link for the Bernoulli distribution in GLMs?
  • What is the MVUE of a normal distribution with a defined variance?
  • GAMs allow for non-linear relationships between each predictor variable and the response.
  • Decreasing the significance level would increase the probability of a Type II error.
  • What is the form of the likelihood function for two independent populations with different density parameters?
  • Forward stepwise selection cannot be used in high-dimensional settings.
  • How many minimal path sets are there for a random graph with n nodes, according to the given material?
  • As lambda increases towards infinity, the ridge penalty term has no effect and the estimates become unconstrained.
  • To model a non-negative response with an unbiased estimate, which error structure and link function combination is most appropriate?
  • Collinearity can exist among three variables even if no single pair shows a high correlation.
  • Under the saturated model, what is the predicted value for each observation?
  • A uniformly MVUE is an estimator such that no other estimator has a smaller variance.
  • Which statement about lasso regression is false?
  • In a regression model that includes an intercept, the sum of residuals is:
  • Mean squared error (MSE) is defined as the expected squared difference between the estimator and the true parameter. Which statement is true?
  • Which expression is the MLE for the exponential distribution when data are censored and truncated?
  • In a Poisson regression model, what is the offset term?
  • In kernel density estimation, what does the cdf indicate we are looking for?
  • Is an absorbing state considered transient or recurrent?
  • For a binary response variable with a continuous explanatory variable, logistic regression is inappropriate.
  • Lasso regression performs variable selection by shrinking some coefficients exactly to zero.
  • Positive recurrence is a class property; if state 2 is positive recurrent, then state 4 must be positive recurrent.
  • In ordinary least squares, the sum of residuals equals which value?
  • In GLMs, the primary consideration for choosing between a Poisson model with a log link and a Gaussian model with an identity link is the distribution of the response variable.
  • True or false: if all states in a finite Markov chain are recurrent, the Markov chain is irreducible.
  • Which Excel expression yields the p-value for a chi-square test?
  • Consistency of an estimator is characterized by the variance converging to zero as the sample size grows.
  • Which expression gives the estimated variance of beta_hat in OLS regression?
  • In simple linear regression, which interval estimates E(Y|X)?
  • For X ~ Uniform(m,n), the conditional distribution X | X > pi_q is Uniform(pi_q, n) provided pi_q lies in (m,n).
  • Which statement about the Neyman-Pearson lemma for simple hypotheses is correct?
  • Using an alternative fitting procedure makes results easier to interpret.
  • Greedy Algorithm B uses which sequence of k values?
  • A consistent estimator is also unbiased.
  • If the number of PLS components equals the number of predictors in OLS, the forecasted values from both methods are what?
  • In curtate life expectancy, ex_(curtate) relates to survival probability p_x and the life expectancy at age x+1 by which expression?
  • A Markov chain that has a limiting distribution is described as which type?
  • For a sample from Uniform(a,b), the expected minimum (k = 1) is a + (b - a)/(n + 1). Which expression correctly represents this?
  • In local regression, increasing the span parameter makes the fit more global rather than local.
  • Ridge regression shrinks the coefficient estimates, which has the benefit of reducing the bias.
  • Which ordering of the degrees of freedom used by the three spline models is correct from most to least, given a linear spline with k knots, a cubic spline with k knots, and a natural cubic spline with k total knots (k minus 2 interior knots)?
  • Which property ensures a Markov chain has a unique stationary distribution and convergence from any starting state?
  • Which statement about a simple linear relationship is true?
  • In PCA, the first principal component is the direction along which the data vary the most.
  • Which method is used to measure the accuracy of a parameter estimate?
  • Overdispersion occurs when the observed variance is larger than the mean for a Poisson model, or when the observed variance is larger than the calculated variance for a binomial model. Which statement best captures overdispersion in common models?
  • In kernel density estimation, after calculating the kernel contributions k_i(x) from each observation, how is the density estimate at x formed?
  • The Cramer-Rao lower bound for the variance of unbiased estimators is given by which expression?
  • What is the MVUE of a binomial distribution?
  • Which distribution is commonly used to model life data due to its flexible hazard function?
  • Among the following kernel options for kernel density estimation, which are symmetric?
  • Which statement about the exponential family and canonical form is true?
  • Dimension reduction reduces the number of predictors by projecting onto a lower-dimensional space.
  • What is a good indication of a probability generating function (PGF)?
  • The variance of error terms doesn't have to be constant.
  • what is the excel equation for solving for the cdf in gaussian kernel estimation?
  • Cluster analysis is typically categorized as which type of learning?
  • In a linear regression model, the leverage for each observation is guaranteed to lie between 1/n and 1.
  • In supervised learning, the variance and the squared bias are inversely related.
  • In Tweedie distributions, data are appropriate when data include zeros and continuous positive values that can be viewed as which distribution?
  • The number of events that occur in disjoint time intervals must be independent is a property of counting processes.
  • If n = 4, how many minimal path sets are there?
  • Which expression is the MSE of an estimator?
  • If n = 4, how many minimal cut sets are there?
  • Which method is used to select the appropriate level of model flexibility?
  • Which distribution uses the canonical link inverse squared?
  • Which term describes a Markov chain that has only one communicating class?
  • If N ~ Poisson(lambda) and S = sum_{i=1}^N X_i where X_i are independent of N with E[X^2] finite, what is Var(S)?
  • Which statistic is used to construct a confidence interval for a single population variance?
  • For a Bernoulli response in a generalized linear model, which set of link functions can be used?
  • Which of the following best describes a typical effect of the L1 penalty in lasso regression?
  • Which statement about ridge regression is true?
  • In a likelihood ratio test, which of the following is a correct statement about the null hypothesis?
  • True or False: Power is the probability of rejecting the null hypothesis, assuming its false.
  • All collinearity problems can be detected by inspection of the correlation matrix.
  • In testing whether a source is significant, the test statistic is the mean square of that source divided by the MSE of the model that has the most predictors.
  • In an exponential family distribution, the sufficient statistic for the natural parameter mu, given observations x1,...,xn, is which of the following?
  • Which statement about deviance is correct?
  • For a sample from an inverse Gaussian distribution, which expression is the MVUE of the mean parameter?
  • A predictor uncorrelated with others has VIF equal to 1.
  • Which expression represents Cov(X,Y) in terms of expectations?
  • Minimal cut sets must have at least one component from each minimal path set.
  • In a normal linear model, the scaled deviance is equal to which of the following?
  • Which kernel density estimator distributes mass uniformly in the neighborhood?
  • Backward stepwise selection cannot be performed on the dataset if n<p.
  • Compared with lasso regression, ridge regression is generally harder to interpret because it retains all predictors in the model.
  • With 3 original variables, what is the maximum number of principal components that can be extracted?
  • In actuarial notation, A_x denotes the present value of a 1-unit death benefit payable at the end of the year of death for a life aged x.
  • R^2 is the fraction of variation in y about the mean of y that's explained by the linear relationship with x.
  • In ridge regression, what happens to squared bias as s increases?
  • What is the primary purpose of fitting a saturated model in generalized linear models?
  • LOOCV requires fitting a model a total of n times.
  • If Y is a complete sufficient statistic for theta and g(Y) is an unbiased estimator of theta, then g(Y) is the MVUE and has the smallest possible variance among all unbiased estimators.
  • If two sampling distributions have the same mean, the one with smaller variance is called what?
  • In a 10-state Markov chain, which property ensures that all states communicate?
  • In regression analysis, does an R-squared value of 0 indicate overfitting?
  • Can a deviance plot be used to visually approximate the MLE of theta?
  • A logit transformation helps in reducing heteroscedasticity.
  • Why do we use w-1 dummy variables for a categorical predictor with w levels?
  • Which of the following statements about the prediction interval and the range it measures is true?
  • If Ti is the time of the ith event, Pr(T2 > 3) represents what probability?
  • Which statement about training set MSE versus test MSE is true?
  • Mallows Cp involves SSE, p, and MSEfull.
  • In Poisson regression, which statement about the variance-mean relationship is true?
  • The MVUE is defined as the unbiased estimator with the minimum variance.
  • In a 5-state Markov chain with two classes {0,1,3} and {2,4}, at least one of the two classes must be recurrent.
  • A Markov chain with two communicating classes is not irreducible.
  • Which statement correctly differentiates homogeneous and non-homogeneous Poisson processes regarding stationary increments?
  • For the acceptance-rejection method, what's the inequality for f(y)/ (c g(y)) relative to a Uniform(0,1) random variable U?
  • Subset selection is used to identify a subset of the predictors and then fit a model using least squares on the reduced set of variables.
  • In simple linear regression, the F-statistic for the model equals the square of the t-statistic for the slope parameter.
  • Bootstrapping can be used to select the appropriate level of model flexibility.
  • Is the density f(y; θ) = θ y for y > θ a member of the exponential family?
  • An estimator is consistent whenever the variance of the estimator approaches zero as the sample size goes to infinity.
  • Under X|θ ~ N(θ, σ^2) and θ ~ N(μ, τ^2), the marginal distribution of X is Normal with mean μ and variance σ^2 + τ^2.
  • How is overdispersion detected in a generalized linear model?
  • Can a plot of Information be used to visually approximate the MLE of theta?
  • ANOVA is a useful approach for analyzing the means of groups of continuous response variables, where the groups are categorical.
  • How is multicollinearity detected using the variance inflation factor (VIF)?
  • What penalty term does ridge regression use?
  • Power parameter in the Tweedie family that corresponds to a normal distribution.
  • When selecting the optimal model, which criterion is preferred?
  • What distribution do you use when calculating a confidence interval for beta coefficients?
  • How do you test for time reversibility of a Markov chain?
  • Power parameter in the Tweedie family that corresponds to a Poisson distribution.
  • For Poisson processes, counts in disjoint intervals are independent.
  • PLS is a subset selection method.
  • Which of the following best describes KNN?
  • In a PCA performed on a data set with 50 observations and 3 independent continuous variables, which statement is true?
  • For modeling a binary outcome such as hospitalization, which distribution and link are most appropriate?
  • The pseudo R-squared is computed as one minus the ratio of the model log-likelihood to the null log-likelihood.
  • For a normal distribution, the deviance is proportional to the residual sum of squares.
  • In ridge regression, test error as s increases?
  • The Cramer-Rao lower bound for the variance of all unbiased estimators of theta equals:
  • Ridge regression shrinks coefficients toward zero and can never set any coefficient exactly to zero.
  • Which statement best describes how lasso influences model sparsity?
  • When forming a confidence interval for the difference between two means with paired observations, which statistic is used?
  • Which distribution has a constant hazard rate?
  • For X ~ Uniform(2,5), conditioning on X > 3 yields X | X > 3 ~ Uniform(3,5).
  • With a cubic spline, what must match at the knot?
  • In Gaussian kernel density estimation, the data point x_i represents what in the kernel sum?
  • Relative to the least squares estimates, shrinkage has the effect of reducing bias.
  • What is the MVUE of a Poisson distribution?
  • In simple linear regression, the statement that the sample correlation between x and y equals the coefficient of determination R^2 is
  • Overfitting causes the training error to underestimate the test error.
  • What is the expected value of the kth order statistic from a Uniform(a,b) distribution?
  • Which statement best describes the difference between homogeneous and non-homogeneous Poisson processes?
  • Which of the following describes the characteristics of a non-homogeneous Poisson random variable?
  • True or false about a smoothing spline model fit to data using the tuning parameter lambda: larger values of lambda result in smoother splines.
  • Backward stepwise selection cannot be used when the dataset has as many predictors as observations (n < p+1).
  • Which sequence of steps correctly calculates the probability of moving from transient state 2 to state 4 in a Markov chain?
  • In the same model, what is Var(X)?
  • If the stationary probability of state 0 in a Markov chain is 0.2, what is the expected return time to state 0?
  • Which term describes a model that achieves adequate predictive performance using the fewest explanatory variables?
  • What is the sum of the leverages across all observations in a linear regression with an intercept and p predictors?
  • A biased estimator can be consistent.
  • If a distribution isn't in canonical form, can there be a natural parameter?
  • In the Tweedie family, p = 1 corresponds to which distribution?
  • PCR can reduce overfitting.
  • The tuning parameter for ridge regression can be selected using cross-validation.
  • Both lasso and ridge regression shrink coefficients toward zero, but lasso can force some coefficients to be exactly zero.
  • Cayley’s formula gives the number of labeled trees on n vertices. What is that count?
  • LOOCV uses n training/testing splits, one for each observation left out.
  • In a scenario where k-out-of-n with k=1 and n=5, the expected lifetime equals theta times the harmonic sum 1 + 1/2 + 1/3 + 1/4 + 1/5. If theta=2, which value is correct?
  • K-fold validation has an advantage over LOOCV in bias reduction.
  • For two independent exponential random variables X and Y, what is E[X | X < Y]?
  • Which interval quantifies the possible range for a future observation Y given X?
  • K-fold validation has variance reduction compared to LOOCV.
  • In marketing analytics, clustering to segment shoppers is best described as which type of learning?
  • In ordinary least squares, the variance of each coefficient estimate is given by which expression?
  • The statement 'The logit link corresponds to a logistic distribution, and the probit link corresponds to a standard normal distribution' is true.
  • In life contingencies notation, Ax denotes the present value of what kind of benefit?
  • For a k-out-of-n system with iid exponential components, the expected lifetime is theta * sum_{i=k}^n 1/i. If theta=2, n=5, k=3, what is the expected lifetime?
  • In exponential family distributions, canonical form implies a(y) equals y.
  • Does the quasi-likelihood approach change the coefficient estimates?
  • The efficiency of theta-hat is the estimator's variance divided by the Rao-Cramer lower bound.
  • Which step is typically performed first when computing expected sojourn times for a Markov chain with transient states?
  • Which method is used when a problem mentions the 'best critical region'?
  • In smoothing spline models, increasing lambda has what effect on the bias-variance tradeoff?
  • True or false: Mallows Cp is an unbiased estimate of the test MSE if its variance is calculated using an unbiased estimate of the variance.
  • With least squares regression on a dataset with n observations, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • Using Cook's distance with a unity threshold, an observation is influential if which condition holds?
  • For an exponential distribution with complete data, what is the maximum likelihood estimator of the mean parameter?
  • In the kernel density estimator context, the contribution of a single kernel centered at x_i to the pdf is proportional to which expression?
  • Given Y claims are made by time t, the unordered times of past events follow what distributions?
  • A limitation of GAMs is that interactions cannot be added to the model.
  • For a cubic spline with one knot, how can we connect the two pieces of the equation to find coefficient estimates?
  • True or False: R-squared is a good measure for model comparison.
  • What is MGF of Y evaluated at t=1?
  • A Markov chain that is irreducible, positive recurrent, and aperiodic is called
  • Ordinal variables are a type of continuous explanatory variable.
  • Which statement best describes an ergodic Markov chain?
  • Which type of qualitative variable has categories with a meaningful order?
  • When forming a confidence interval for the difference of two proportions, which statistic is used?
  • n^(n-2) is associated with the count of which standard combinatorial object?
  • The sum of leverages across observations equals p+1.
  • The squared bias increases as the method's flexibility decreases.
  • Which Poisson process type has stationary increments?
  • Removing any component from a minimal path set will break the guarantee.
  • Poisson regression models assume exposure is constant when modeling the rate at which events occur.
  • In ridge regression, training error as s increases?
  • In ridge regression, coefficients shrink toward zero but are typically not exactly zero.
  • What could be added to a linear regression model to avoid multicollinearity?
  • Power parameter in the Tweedie family that corresponds to an inverse-Gaussian distribution.
  • PLS is a dimension reduction method.
  • K-fold cross validation requires fitting a model for a total of k times.
  • The law of total variance states Var(X) = E[Var(X|θ)] + Var[E(X|θ)].
  • What is the test statistic for a likelihood ratio test?
  • What is the only continuous distribution listed with a finite support?
  • For a Gamma distribution with known shape parameter alpha and i.i.d. samples, what is the MVUE of the scale parameter beta?
  • Ridge regression coefficients are not scale equivariant.
  • Does implementing the quasi-likelihood approach to a GLM change the coefficient estimates?
  • Residual sum of squares is a suitable metric for selecting the best model among models with different numbers of predictors.
  • In a Poisson process, increments over disjoint time intervals are what property?
  • An annuity-due issued to a 40-year-old that pays 1 each year until death or age 60, whichever comes first, has an actuarial present value given by which expression?
  • In Greedy Algorithm A, after selecting the lowest-cost assignment, what is done next?
  • Cp, AIC, BIC, and adjusted R-squared are used to adjust which type of error when evaluating models for model size?
  • A GLM with the same distribution and link function as the model of interest that has the max number of parameters that can be estimated; assists in assessing model adequacy is called what?
  • Greedy Algorithm B for optimization problems is described as which approach?
  • Given that the corresponding coefficient estimates for all models with two explanatory variables are the same, which link function will produce a prediction for an observation that is always the greatest?
  • A minimal path set is a minimal set of components whose functioning guarantees the functioning of the system.
  • When forming a confidence interval for the ratio of two variances, what kind of statistic is used?
  • Ridge regression uses an L2 penalty, while lasso regression uses an L1 penalty.
  • Forward stepwise selection requires fitting 1 + (1/2)(p*(p+1)) models.
  • Which type of qualitative variable has categories without a meaningful order?
  • As model flexibility increases, the test MSE monotonically decreases.
  • If we want the expected number of time periods a chain is spent in state j given it started in state i, what are the steps for calculating this?
  • In least squares regression, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • For a sample from an inverse Gaussian distribution, the MVUE of the mean parameter is:
  • The maximum possible value of the standard error of the population variance is achieved under which condition?
  • Removing high leverage points in a linear regression model primarily affects which aspect of the fitted model?
  • In the Tweedie family, p = 0 corresponds to which distribution?
  • As more variables are added to a linear model, the training MSE generally does what?
  • What penalty term does lasso regression use?
  • Power parameter in the Tweedie family that corresponds to a compound Poisson-Gamma distribution (1<p<2).
  • Is a cluster analysis an example of supervised or unsupervised learning?
  • A counting process possesses independent increments if the number of events between s and t is independent of the number between t and t+u for all u>0.
  • In a Poisson process, the waiting time until the first event is exponentially distributed with a parameter theta equal to what value in terms of lambda?
  • In ridge regression, irreducible error as s increases?
  • Using quasi-likelihood, what must be assumed about the relationship between the mean and the variance?
  • True or false: Ridge regression is less flexible than OLS and thus results in an improved prediction accuracy when its increase in squared bias is less than its decrease in variance.
  • In PCR, it is common to use only the first few principal components to predict the response.
  • In the transient-state fundamental matrix S = (I - PT)^{-1}, what does the entry S_ij represent?
  • In simple and multiple linear regression, the maximum likelihood estimator for the regression coefficients coincides with the ordinary least squares estimates when residuals are normally distributed.
  • Shrinkage fits a model involving a subset of predictors with the estimated coefficients shrunken towards zero.
  • A high leverage point is an outlier.
  • A logit model applies when the explanatory variables are both continuous and categorical.
  • The deviance is useful for testing the significance of explanatory variables in nested models.
  • In the context of Poisson residuals, the Pearson residual is defined using which denominator?
  • Which of the following is generally considered supervised learning?
  • When scaling X ~ Lognormal(mu, sigma^2) by a positive constant c, which distribution describes cX?
  • The statement that the proportion of variance explained by an additional principal component increases as more PCs are added is true or false?
  • Regression through the origin occurs when the intercept term in the linear equation linking the explanatory variables to the dependent variable is zero or is left out of the equation. Which of the following is true?
  • Which statement about ridge regression is true?
  • For theta=0.5, n=4, k=2, what is the numeric value of E when E = theta * sum_{i=k}^n 1/i?
  • PCA provides low-dimensional linear surfaces that are closest to the observations.
  • For an exponential random variable X, what is E[X | X > a]?
  • When forming a confidence interval for a mean with unknown variance, which statistic is used in the calculation?
  • A uniformly MVUE is defined as an estimator such that no other estimator has a smaller variance.
  • In the Poisson context, which statement describes overdispersion?
  • Dimension reduction involves projecting the p predictors into an M-dimensional subspace by computing M distinct linear combinations of the variables and utilizing them as predictors for fitting a linear regression model.
  • In linear regression with p predictors and an intercept, the sum of leverages equals which of the following?
  • Can a plot of the score function be used to visually approximate the MLE of theta?
  • In the context of regression, the statement 'the sample correlation between x and y is equal to the coefficient of determination' is
  • Adjusted R-squared equal to 1 indicates overfitting.
  • Which statement about the relationship between Mallows Cp and AIC is most consistent with the material?
  • In life-table notation, p_x denotes the probability of surviving from age x to age x+1.
  • Which statement best describes how model flexibility affects variance and bias?
  • Residual plots are a useful graphical tool for identifying non-linearity.
  • According to the Central Limit Theorem, what is the limiting distribution of the sampling distribution of the sample mean?
  • If an explanatory variable is uncorrelated with all other explanatory variables, the corresponding variance inflation factor would be 1.
  • A logit model gives numerical results that are quite similar to those given by the probit model.
  • Only w-1 dummy variables are needed to represent w classes of a categorical predictor.
  • Which statement best indicates a variable is statistically significant in a standard hypothesis test?
  • Which model uses more parameters for the same number of knots: a linear spline with k knots or a natural cubic spline with k total knots?
  • What's the formula for Var(X+Y)?
  • Loadings for the first direction are proportional to covariances between the response and each standardized predictor.
  • In kernel density estimation, what does the pdf indicate we are looking for?
  • PLS identifies new features in a supervised way by relating them to the target variable.
  • Homoscedasticity occurs when what condition holds?
  • True or false: For the same total number of knots k, the natural cubic spline uses fewer degrees of freedom than the cubic spline.
  • In regression with Gaussian errors, a large value of Mallows' Cp indicates a model with a low test error.
  • Which kernel density estimation fact is true?
  • A biased estimator can only be inconsistent.
  • Ridge regression coefficients are not scale equivariant.
  • Ordinary least squares estimators are inherently unbiased.
  • For X ~ Exp(λx) and Y ~ Exp(λy) with 1 < X < Y, what is E[X | 1 < X < Y]?
Subscribe

Get the latest from Examzify

You can unsubscribe at any time. Read our privacy policy