DEV Community

Cover image for PennyWise Economic Decisions: Benchmarking Frontier LLMs on Financial Reasoning Efficiency
artespraticas
artespraticas

Posted on

PennyWise Economic Decisions: Benchmarking Frontier LLMs on Financial Reasoning Efficiency

Kaggle Benchmarking Challenge Submission

What I Benchmarked

I created the PennyWise Economic Decisions benchmark task to evaluate how effectively modern large language models handle complex, multi-variable financial reasoning and economic scenario classifications.

The task exposes models to 8 distinct economic scenarios designed to test practical fiscal decision-making. The metric measures not only accuracy in choosing the correct financial option but also the model's efficiency—calculating token spend versus minimum necessary spend to reveal which models deliver the best ROI (Return on Investment) for financial automation pipelines.

Models Tested

I evaluated an expansive, diverse cross-section of frontier and miniature models to see how scale impacts financial precision and budget allocation:

  • Google Gemini 3.5 Flash & 2.5 Flash (To observe speed and lightweight economic parsing capabilities)
  • OpenAI GPT-5.5 & GPT-5.4 mini (To test cutting-edge reasoning against optimized smaller variants)
  • Anthropic Claude Sonnet 4.6 & Claude Opus 4.6 (To evaluate strict rule-following and structured output constraints)
  • xAI Grok 4.6
  • DeepSeek R1 (To evaluate how specialized reasoning models parse financial chains of thought)
  • Qwen3 235B (To check open-weights performance at scale)

Findings

The results were incredibly revealing, specifically highlighting the prowess of highly optimized modern architectures:

  • Perfect Scores at Fraction of the Cost: Gemini 3.5 Flash performed flawlessly across all metrics. It handled the financial logic effortlessly, yielding a PennyWise Score of 100.0 and a 1.0 average efficiency rating.
  • Zero Waste Efficiency: The most surprising insight was the absolute optimization of the token usage. Gemini 3.5 Flash spent exactly $0.035 to finish the entire suite, recording a Total Waste of $0, matching the exact theoretical minimum necessary spend.
  • What's Next: For future iterations of this benchmark, I intend to introduce highly adversarial data inputs, such as conflicting market indicators and complex JSON schema constraints, to push the reasoning ceilings of these frontier models even further.

My Benchmark

You can explore the full implementation details, model scoring logs, and real-time execution leaderboard directly on Kaggle here:
👉 PennyWise Economic Decisions Task Page

Note: This entire benchmark suite was successfully built, configured, and executed using a cloud-based Kaggle Notebook directly inside a Chrome browser tab on an iPad!

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to