Abstract

This study explores the feasibility of teaching Large Language Models (LLMs) to predict future events through a novel retrospective evaluation framework. We fine-tune Llama-3-8B models on real prediction-market data and employ an autonomous, time-sensitive news retrieval system to make probabilistic forecasts.

Key Contributions

  1. Novel Pipeline: Introduced an automated pipeline that leverages LLMs and web search to synthesize high-quality, verifiable forecasting datasets
  2. Fine-tuning Framework: Established a robust framework that aligns language models with probabilistic forecasting tasks
  3. Autonomous Retrieval: Implemented a news retrieval architecture that aggregates and synthesizes relevant context to ground predictions

Results

Our results demonstrate that fine-tuning significantly reduces Brier scores compared to zero-shot approaches, indicating improved calibration. Notably, we found no significant performance difference between domain-specific and general-purpose models, suggesting that fine-tuning enhances generalized reasoning capabilities rather than relying on domain-specific memorization.

Authors

Andy Kaboski, Gabriel López-Asiaín, Karsten Kropp

University of Chicago