Support tickets, judged right.
900 rule-generated tickets, three questions each. The reference system scores 75.1% on the same set — the gap is published below, not buried.
- SemIf-144 accuracy
- 65.3%
- Queue routing (4-way)
- 93.0%
- Priority scoring
- 87.3%
- Median, hosted
- 337 ms
Benchmarked on 900 unseen tickets
Independent OOD bench (source): 300 queue-routing choices, 300 mood booleans, 300 priority scores. Reflex-1 ran the full suite in September 2026, one judgment per question. Bars show accuracy; ECE figures are listed below.
SemIf-authored144: Reflex-1 65.3% accuracy (94/144 valid responses) — chance is 33.3%. OOD runs: single-judgment inference, temp 0. Reflex-1 learned related ticket templates, so treat the gap as improvement on this task family, not proof on unseen domains.
Run a decision here
Send a live POST /v1/classifier to the hosted worker. Keys stay in this browser.
Mint a key. Keep it somewhere safe.
Keys are free during the open demo. One per signup, shown once, stored hashed.
Two ways to watch it decide.
Reflex Pilot
A city car driven entirely by classifier calls — every steering choice streams from this API at ~300 ms. It drove a full 735 m route with zero contacts.
735 m · 0 contacts · J to engage
Play driving2048
One board, one classifier call per move. The model reached tile 64 in 74 moves with zero invalid answers. Play manually with arrows, or let it play.
74 moves · tile 64 · arrows to play
Play 2048Tetris
One piece, one classifier call, one placement chosen from a shortlist. The model cleared 5 lines across a 49-piece game in a real browser.
49 pieces · 5 lines · Space to drop
Play TetrisSimulation, physics, game rules, and artwork © their authors; credits ship in each game and in the project’s attribution files.