How Magicdoor Tested Auto Routing—and What the Test Cannot Prove
Direct answer: Magicdoor tested router decisions, not generated-answer superiority. In a historical discovery exercise, 164 unique synthetic scenarios produced 324 live TypeSafe Jev requests including repeats. A fresh final audit matched the author's intended route labels on 28 of 28 cases. That is useful evidence about an earlier decision boundary—not a claim that Experimental Auto gives better answers, costs less overall, or works perfectly on live traffic.
Fact-checked October 2, 2026. This page owns the study design and limitations. The exact current Auto policy owns live behavior; the evaluation rubric explains the additional proof a buyer should seek.
What was measured
| Phase | Unique synthetic cases | Purpose |
|---|---|---|
| Development | 81 | Compare question wording and gates, including previous cases and fresh additions |
| First holdout | 31 | Test frozen preferred candidate and comparators on unseen scenarios |
| Confirmation | 24 | Probe newly discovered challenge and creative-work distinctions |
| Final audit | 28 | Test the combined question at a frozen 0.30 threshold on fresh cases |
These phases total 164 distinct scenarios. The 324 live Jev requests also include repeated boundary calls; they are not 324 independent people or 324 separate answers. Eighteen candidate wording/gating combinations were compared (15 initial and three exploratory threshold additions). Two repeated-boundary requests failed or exceeded a 2.5-second deadline; they were not silently counted as successful classifications.
The evaluation asked Jev typed questions against example conversation states and compared the selected route to a product-policy label assigned by the study author. Examples included routine rewriting, disputed recommendations, code writing versus code execution, file-capability flags and a few non-English turns. No customer conversations or real attachment contents were sent for this study. The original label set was about the earlier candidate policy, not a prospective benchmark of today's four-model pool.
Why the final audit matters—but is limited
The combined quality question at threshold 0.30 matched 28/28 intended labels in a final audit after the choice was frozen. The aggregate agreement on all 164 cases is partly in-sample: earlier phases helped choose that threshold. The initial preferred candidate was frozen before the first holdout, but switching to the combined question after confirmation and repetition is itself model selection. The final audit reduced this particular selection bias; it cannot eliminate author-label bias, related examples or narrow language coverage.
A matching route label is not a correct answer. The study did not generate final responses from competing models, blind-review correctness or constraints, compare a direct-model answer baseline, or measure production p95, total billed cost or conversion. Repeated calls test boundary stability under their own conditions, not a live service-level objective. The discovery request batches used different timeouts from the live router, so their timings must not be treated as today's latency.
What changed since that study
The discovery report describes an earlier combined-gate experiment and old model options. The live application now uses Experimental Auto, quality-first-v7, and a pool of DeepSeek V4.1 Flash, GPT-6 Luna, GPT-6.1 Sol, and Claude Opus 5.5, with a possible early DeepSeek-to-Opus handoff and capability-aware fallback. The current server policy and registry—not the discovery report's historical recommendation—govern each message. TypeSafe AI created Jev; this test is Magicdoor's use of its API, not TypeSafe endorsement.
For a trust decision, follow the full router evaluation rubric: blind answer comparisons against a direct baseline, cost including retries, p95 latency and safety/capability failures on representative workload slices. Check exactly what text Auto shares with Jev if the data boundary matters to you. Manual model selection bypasses Jev routing.
FAQ
Did Magicdoor test 164 real customer conversations?
No. The 164 cases were synthetic scenarios with author-assigned expected route labels. No customer conversations or actual attachment contents were sent for that discovery study.
Does 28 out of 28 final-audit matches mean Auto always chooses the best answer?
No. Those 28 matches compare routes with labels on a fresh synthetic audit. The study did not generate and blindly compare final answers or establish production quality, savings, or reliability.
Is the original study a live benchmark of the current four-model Auto pool?
No. The discovery study informed an earlier routing threshold and wording. Current Experimental Auto uses quality-first-v7 and a four-model pool; the old label counts do not benchmark that current answer policy.
Sources
Accessed October 2, 2026.
- Magicdoor's recorded
docs/typesafe-routing-discovery.mdand its reproducible experiment artifacts underscripts/experiments/typesafe-discovery/for phase counts, failures and caveats; historical rather than a current runtime specification. - Current
typesafeRoutingPolicy.ts,typesafeRouter.tsandmodelRegistry.tsfor current policy and answer-model pool. - TypeSafe AI: State for how typed questions evaluate supplied state.
Related Resources
What Magicdoor Auto Shares With Jev (and What It Does Not)
See the bounded text state Magicdoor's Experimental Auto sends to TypeSafe AI's Jev, which parts are excluded, and why exclusion is not complete redaction.
ChatGPT Auto vs Magicdoor Auto: What Each One Actually Selects
Compare ChatGPT's current automatic reasoning switch with Magicdoor Experimental Auto across models, providers, controls, context, billing, and limitations.
OpenRouter Auto vs Magicdoor Auto: API Router or AI Workspace?
Compare OpenRouter Auto with Magicdoor Experimental Auto, including routing method, model control, billing, context, fallbacks, and who each product serves.
How Magicdoor Auto Chooses DeepSeek, Luna, Sol, or Opus
See the exact Experimental Auto policy Magicdoor uses to choose DeepSeek V4.1 Flash, GPT-6 Luna, GPT-6.1 Sol, or Claude Opus 5.5 for each chat message.