This study explores the feasibility of teaching Large Language Models (LLMs) to predict future events through a novel retrospective evaluation framework. We fine-tune Llama-3-8B models on real prediction-market data and employ an autonomous, time-sensitive news retrieval system to make probabilistic forecasts.
Our results demonstrate that fine-tuning significantly reduces Brier scores compared to zero-shot approaches, indicating improved calibration. Notably, we found no significant performance difference between domain-specific and general-purpose models, suggesting that fine-tuning enhances generalized reasoning capabilities rather than relying on domain-specific memorization.
Andy Kaboski, Gabriel López-Asiaín, Karsten Kropp
University of Chicago