Five Hypotheses for Why LLMs Fail at Tabular Data.
A systematic study rules out four plausible explanations for why LLMs underperform classical ML on tabular classification, and finds the real culprit: accuracy degrades with feature count in a way no noise-corrupted classical model reproduces, and the model's own explanations don't match what it computed.