
3 AI models call live Polymarket markets daily. Public Brier scores. Every divergence logged. Nothing edited. Day 8 of 365.
Loading comments…
Visit Website
emberfyi.com
Project Info
Product Keywords
Achievement
Maker
Ember is a live AI prediction experiment that runs three independent AI models—Claude (Anthropic), Grok (xAI), and Gemini (Google)—against real Polymarket markets every day. Each model makes a probability call before the market resolves, and every call is logged, scored, and published publicly. The project runs on a strict 365-day methodology with locked rules, public Brier scores, and zero editing of predictions after they're made.
Ember runs Claude, Grok, and Gemini on the same market questions daily. Each model uses a different reasoning approach—Claude synthesizes prediction markets and forecaster communities, Grok reads live X sentiment, and Gemini grounds calls in search results. When models disagree, that divergence becomes the signal.
Every resolved call is scored using the Brier score, a standard accuracy metric for probabilistic predictions. The scoreboard is fully public and updated daily. After 8 days live, Ember holds a Brier of 0.0365 against a crowd score of 0.0356 across 157 resolved calls.
Predictions are never edited after being made. When infrastructure issues affect data integrity, Ember publishes public correction notices—9 so far. The methodology stays locked for the full 365-day arc, with a Year 2 evaluation beginning at Day 300.
When a model's probability diverges significantly from the crowd market price, Ember flags it as a high-conviction signal. These alerts are available in real-time to subscribers and highlight potential mispricings before they resolve.
Three AI models call live markets daily, independently, before the outcome. Nothing edited. Every divergence logged and scored.
Ember's edge is its commitment to radical transparency in AI prediction. Unlike black-box forecasting systems, Ember publishes its full record—including correction notices for infrastructure issues—and never retroactively changes a call. The 365-day experiment forces the models to build a track record over time, making the Brier scores meaningful for evaluating whether AI can actually predict its own trajectory.
Other tools you might consider
You're interested in AI prediction accuracy as a measurable, public experiment rather than marketing claims. Traders looking for systematic divergence signals, developers building on prediction market data, and researchers studying AI forecasting will all find value in Ember's open methodology and live scoreboard. The $29/month Founding Arena tier gives access to live calls before public release, while the API starts at $299/month for institutional integration.