Tijori
2026-09-08 15:08
Buoyant
2026-09-07 (model valuation date)
41 lenders598 syncs ok2 failed

Regression Lab & Driver Discovery

§26/§27 — what has actually explained Indian bank returns and multiples in this dataset.

Correlation does not establish causality. The panel is annual FY2017–FY2026 (≤10 observations per bank, 386 bank-years). Stock returns are Derived from P/B × book value plus dividend yield, because Tijori has no price time series. Tijori also provides no release dates, so fundamentals are matched to the reporting period, not the date they became public — some look-ahead bias remains. Read the t-statistics accordingly.
Specification. Dependent variable is annual total return per share, winsorised at 1/99. Panel is 242 bank-years across 40 banks. Every estimator below is run on the same data so the columns are comparable, and standard errors are reported under four different assumptions because the choice changes the conclusion.

The result that matters most is at the bottom. A random 5-fold split gives an R² of 0.26 — but that leaks future information into the training set. Blocking by year, which is the only honest test, gives -0.08. The linear model does not predict out of sample. It describes what happened; it does not forecast.

Pooled OLS — the same coefficient under four standard-error assumptions

Two-way clustering by bank and year is the defensible one: returns are correlated within a bank over time and across banks within a year.

VariableCoefficientSE (OLS)SE (HC1)SE (2-way cluster)tp
Book value / share growth1.28820.1960.2170.1817.120.000
Starting P/B-19.58372.8262.6114.591-4.270.000
Δ ROE-0.20580.4850.4280.701-0.290.769
Loan growth0.63410.2670.2980.2312.740.006
Δ credit cost-4.19104.1353.8475.549-0.760.450

adj R² 0.277 · AIC 2572 · BIC 2593. Clustered SEs are up to 60% wider than naive OLS on starting P/B — ignoring the panel structure would have overstated significance.

Estimator comparison — t-statistics

A finding that survives every estimator is real. One that appears in only some is an artefact of the assumption.

VariablePooled (cl)FE entityFE timeFE two-wayFama-MacBethDriscoll-KraayVerdict
Book value / share growth7.12*4.90*8.31*10.92*0.259.97*robust (5/6)
Starting P/B-4.27*-3.29*-3.00*-1.92-2.23*-2.83*robust (5/6)
Δ ROE-0.29-0.57-0.07-0.581.41-0.07not significant (0/6)
Loan growth2.74*1.550.820.15-0.380.78fragile (1/6)
Δ credit cost-0.76-0.48-0.02-0.20-0.09-0.02not significant (0/6)

Read this table before any other. Starting P/B is the only variable significant under every estimator including Fama-MacBeth — buying expensive reliably hurts. Book-value growth is strong in the panel but collapses under Fama-MacBeth (t 0.25), which means the cross-sectional relationship is unstable year to year: it works on average across the decade, not in any given year. Loan growth is significant pooled and dies under fixed effects — a between-bank difference, not something a bank earns by accelerating. Δ ROE is not significant anywhere.

Specification tests

Whether the model is entitled to the inference it reports.

Breusch-Pagan (heteroskedasticity)0.1728no evidence of heteroskedasticity
White test0.2534
Jarque-Bera (normality)0.0000residuals are not normal — rely on the clustered/bootstrap inference, not t-tables
Durbin-Watson2.33~2 means no first-order serial correlation
Ramsey RESET (functional form)0.0007rejects linearity — see model selection below
Condition number47.2below 30 is comfortable
Hausman (FE vs RE)0.0006reject random effects — bank-specific effects correlate with the regressors, so fixed effects is the consistent estimator
VIF: Book value / share growth 1.12 · Starting P/B 1.13 · Δ ROE 2.65 · Loan growth 1.19 · Δ credit cost 2.64 — all VIF below 5 — no collinearity problem

Quantile regression

Do the drivers differ in a bad year versus a good one? Coefficients at the 10th, 50th and 90th percentile of returns.

Variableτ=0.10τ=0.50τ=0.90
Book value / share growth1.1961.2671.215
Starting P/B-15.952-15.312-27.693
Δ ROE-0.8960.043-0.425
Loan growth0.8720.6290.918
Δ credit cost-9.749-1.168-5.205

Starting P/B punishes hardest in the best years (τ=0.90 coefficient roughly double the median): in a strong tape, the expensive names do not participate. Book-value growth is stable across the whole distribution — it pays in good years and bad, which is what you want from a core driver.

Model selection and functional form

RESET rejected linearity, so non-linear terms are tested rather than assumed. Ranked by BIC, with honest blocked-by-year out-of-sample R².

Specificationnadj R²AICBICRESET pOOS R² (blocked)
linear (base)2420.277257225930.001-0.079
+ starting P/B squared2420.312256125860.000-0.007
+ log starting P/B2420.354254525660.0000.098
+ BV growth x starting P/B2420.274257425990.000-0.088
+ both non-linear terms2420.316256125890.000-0.005

Lowest BIC: '+ log starting P/B'. RESET on the linear base rejects at p=0.0007, so the linear form is misspecified; the non-linear terms are tested rather than assumed. Note that none of them rescues out-of-sample performance.

The winner is economically sensible rather than a fitting artefact: valuation should act multiplicatively, not linearly. Going from 1x book to 2x is a doubling; going from 4x to 5x is a quarter. Taking logs says the same thing the arithmetic does. It is also the only specification with a positive out-of-sample R², and even that is 0.098 — real, but small. Use this to understand what has driven returns, not to forecast them.