Support tickets, judged right.
900 rule-generated tickets, three questions each. The reference system scores 76.3% on the same set — the gap is published below, not buried.
- SemIf-144 accuracy
- 78.5%
- Queue routing (4-way)
- 95.3%
- Priority scoring
- 92.0%
- Median, hosted
- 317 ms
Benchmarked on unseen tickets, intents, and entailment
Independent OOD bench (source): 300 queue-routing choices, 300 mood booleans, 300 priority scores — plus SemIf-144 and Banking77-1200 (20-way intents, reserved test split). Reflex-1 ran the full suite in September 2026, one judgment per question. Jev re-measured 2026-09-24 via OpenRouter (pinned typesafe/jev-1.13-20260917). Bars show accuracy.
SemIf-144: Reflex-1 78.5% (113/144) vs Jev 97.2% (140/144) — SemIf is public and likely in Jev's training mix; chance is 33.3%. Banking77-1200 (20-way intents, chance 5%): Reflex-1 89.7% vs Jev 88.5% on the reserved test split. OOD runs: single-judgment inference, temp 0. Reflex-1 learned related ticket templates, so treat the gap as improvement on this task family, not proof on unseen domains.
Mint a key. Keep it somewhere safe.
Keys are free during the open demo. One per signup, shown once, stored hashed.
Two ways to watch it decide.
Reflex Pilot
A city car driven entirely by classifier calls — every steering choice streams from this API at ~300 ms. It drove a full 735 m route with zero contacts.
735 m · 0 contacts · J to engage
Play driving2048
One board, one classifier call per move. The model reached tile 64 in 74 moves with zero invalid answers. Play manually with arrows, or let it play.
74 moves · tile 64 · arrows to play
Play 2048Tetris
One piece, one classifier call, one placement chosen from a shortlist. The model cleared 5 lines across a 49-piece game in a real browser.
49 pieces · 5 lines · Space to drop
Play TetrisSimulation, physics, game rules, and artwork © their authors; credits ship in each game and in the project’s attribution files.