Regression Lab & Driver Discovery
§26/§27 — what has actually explained Indian bank returns and multiples in this dataset.
The result that matters most is at the bottom. A random 5-fold split gives an R² of 0.26 — but that leaks future information into the training set. Blocking by year, which is the only honest test, gives -0.08. The linear model does not predict out of sample. It describes what happened; it does not forecast.
Pooled OLS — the same coefficient under four standard-error assumptions
Two-way clustering by bank and year is the defensible one: returns are correlated within a bank over time and across banks within a year.
| Variable | Coefficient | SE (OLS) | SE (HC1) | SE (2-way cluster) | t | p |
|---|---|---|---|---|---|---|
| Book value / share growth | 1.2882 | 0.196 | 0.217 | 0.181 | 7.12 | 0.000 |
| Starting P/B | -19.5837 | 2.826 | 2.611 | 4.591 | -4.27 | 0.000 |
| Δ ROE | -0.2058 | 0.485 | 0.428 | 0.701 | -0.29 | 0.769 |
| Loan growth | 0.6341 | 0.267 | 0.298 | 0.231 | 2.74 | 0.006 |
| Δ credit cost | -4.1910 | 4.135 | 3.847 | 5.549 | -0.76 | 0.450 |
adj R² 0.277 · AIC 2572 · BIC 2593. Clustered SEs are up to 60% wider than naive OLS on starting P/B — ignoring the panel structure would have overstated significance.
Estimator comparison — t-statistics
A finding that survives every estimator is real. One that appears in only some is an artefact of the assumption.
| Variable | Pooled (cl) | FE entity | FE time | FE two-way | Fama-MacBeth | Driscoll-Kraay | Verdict |
|---|---|---|---|---|---|---|---|
| Book value / share growth | 7.12* | 4.90* | 8.31* | 10.92* | 0.25 | 9.97* | robust (5/6) |
| Starting P/B | -4.27* | -3.29* | -3.00* | -1.92 | -2.23* | -2.83* | robust (5/6) |
| Δ ROE | -0.29 | -0.57 | -0.07 | -0.58 | 1.41 | -0.07 | not significant (0/6) |
| Loan growth | 2.74* | 1.55 | 0.82 | 0.15 | -0.38 | 0.78 | fragile (1/6) |
| Δ credit cost | -0.76 | -0.48 | -0.02 | -0.20 | -0.09 | -0.02 | not significant (0/6) |
Read this table before any other. Starting P/B is the only variable significant under every estimator including Fama-MacBeth — buying expensive reliably hurts. Book-value growth is strong in the panel but collapses under Fama-MacBeth (t 0.25), which means the cross-sectional relationship is unstable year to year: it works on average across the decade, not in any given year. Loan growth is significant pooled and dies under fixed effects — a between-bank difference, not something a bank earns by accelerating. Δ ROE is not significant anywhere.
Specification tests
Whether the model is entitled to the inference it reports.
| Breusch-Pagan (heteroskedasticity) | 0.1728 | no evidence of heteroskedasticity |
| White test | 0.2534 | |
| Jarque-Bera (normality) | 0.0000 | residuals are not normal — rely on the clustered/bootstrap inference, not t-tables |
| Durbin-Watson | 2.33 | ~2 means no first-order serial correlation |
| Ramsey RESET (functional form) | 0.0007 | rejects linearity — see model selection below |
| Condition number | 47.2 | below 30 is comfortable |
| Hausman (FE vs RE) | 0.0006 | reject random effects — bank-specific effects correlate with the regressors, so fixed effects is the consistent estimator |
Quantile regression
Do the drivers differ in a bad year versus a good one? Coefficients at the 10th, 50th and 90th percentile of returns.
| Variable | τ=0.10 | τ=0.50 | τ=0.90 |
|---|---|---|---|
| Book value / share growth | 1.196 | 1.267 | 1.215 |
| Starting P/B | -15.952 | -15.312 | -27.693 |
| Δ ROE | -0.896 | 0.043 | -0.425 |
| Loan growth | 0.872 | 0.629 | 0.918 |
| Δ credit cost | -9.749 | -1.168 | -5.205 |
Starting P/B punishes hardest in the best years (τ=0.90 coefficient roughly double the median): in a strong tape, the expensive names do not participate. Book-value growth is stable across the whole distribution — it pays in good years and bad, which is what you want from a core driver.
Model selection and functional form
RESET rejected linearity, so non-linear terms are tested rather than assumed. Ranked by BIC, with honest blocked-by-year out-of-sample R².
| Specification | n | adj R² | AIC | BIC | RESET p | OOS R² (blocked) |
|---|---|---|---|---|---|---|
| linear (base) | 242 | 0.277 | 2572 | 2593 | 0.001 | -0.079 |
| + starting P/B squared | 242 | 0.312 | 2561 | 2586 | 0.000 | -0.007 |
| + log starting P/B | 242 | 0.354 | 2545 | 2566 | 0.000 | 0.098 |
| + BV growth x starting P/B | 242 | 0.274 | 2574 | 2599 | 0.000 | -0.088 |
| + both non-linear terms | 242 | 0.316 | 2561 | 2589 | 0.000 | -0.005 |
Lowest BIC: '+ log starting P/B'. RESET on the linear base rejects at p=0.0007, so the linear form is misspecified; the non-linear terms are tested rather than assumed. Note that none of them rescues out-of-sample performance.
The winner is economically sensible rather than a fitting artefact: valuation should act multiplicatively, not linearly. Going from 1x book to 2x is a doubling; going from 4x to 5x is a quarter. Taking logs says the same thing the arithmetic does. It is also the only specification with a positive out-of-sample R², and even that is 0.098 — real, but small. Use this to understand what has driven returns, not to forecast them.