GATEverse Practice, past papers & mock tests

Formula Vault · ML

Machine Learning

🔖
Sheet 1

Regression, Classification & Model Evaluation

7 formulas
Linear Regression (Normal Equation) ☆
\[ \hat{\beta} = (X^TX)^{-1}X^Ty \]
Ridge Regression ☆
\[ \hat{\beta} = (X^TX + \lambda I)^{-1}X^Ty \]
shrinks coefficients, never to exactly 0
Logistic (Sigmoid) Function ☆
\[ \sigma(z) = \frac{1}{1+e^{-z}} \]
Cross-Entropy (Log) Loss ☆
\[ L = -\big[y\log(\hat{y}) + (1-y)\log(1-\hat{y})\big] \]
Naive Bayes Classifier ☆
\[ P(y \mid x) \propto P(y)\prod_i P(x_i \mid y) \]
Precision, Recall & F1 ☆
\[ P=\frac{TP}{TP+FP} \quad R=\frac{TP}{TP+FN} \quad F1=\frac{2PR}{P+R} \]
Bias-Variance Decomposition ☆
\[ \text{Error} = \text{Bias}^2 + \text{Variance} + \text{Irreducible Error} \]

Can you recall the linear regression (normal equation) formula?

Reveal formula
\[ \hat{\beta} = (X^TX)^{-1}X^Ty \]
Accuracy is misleading on imbalanced data
A classifier that always predicts the majority class can still score high accuracy on an imbalanced dataset -- precision, recall and F1 are the metrics that expose this.
High bias = underfitting, high variance = overfitting
A frequently swapped pair: high bias means the model is too simple (underfits both train and test); high variance means it's too sensitive to the training data (overfits -- great on train, poor on test).
Ridge shrinks coefficients, it doesn't zero them out
Ridge regression's L2 penalty shrinks every coefficient toward zero but never sets one exactly to zero -- it doesn't perform feature selection the way an L1 penalty would.
Sheet 2

Unsupervised Learning & Neural Networks

6 formulas
k-Means Objective ☆
\[ \min \sum_{k} \sum_{x_i \in C_k} \lVert x_i - \mu_k \rVert^2 \]
Single-Linkage Distance ☆
\[ d(A,B) = \min_{a \in A,\, b \in B} d(a,b) \]
Covariance Matrix ☆
\[ \text{Cov}(X) = \frac{1}{n}(X-\bar{X})^T(X-\bar{X}) \]
PCA: Variance Explained ☆
\[ \text{explained}_i = \frac{\lambda_i}{\sum_j \lambda_j} \]
λ_i = eigenvalues of the covariance matrix
Perceptron Update Rule ☆
\[ w \leftarrow w + \eta (y - \hat{y})\,x \]
Gradient Descent Update ☆
\[ \theta \leftarrow \theta - \eta \nabla L(\theta) \]

Can you recall the k-means objective formula?

Reveal formula
\[ \min \sum_{k} \sum_{x_i \in C_k} \lVert x_i - \mu_k \rVert^2 \]
k-means needs k chosen in advance
Unlike hierarchical clustering, k-means requires the number of clusters k fixed up front, and its result is sensitive to the initial centroid placement -- different runs can converge to different local optima.
PCA maximizes variance, not class separability
PCA picks directions of maximum VARIANCE in the data -- these aren't necessarily the directions that best separate classes for a downstream classifier (that's what LDA is for instead).
A larger k in k-NN isn't automatically better
Too small a k overfits to noise; too large a k oversmooths and underfits -- k is a bias-variance trade-off knob, not a 'bigger is safer' parameter.