Models / Sansi
Sansi
The same DeepSeek V4 Flash: draft first, review, and revise only if a defect is found, with mechanical gates on top. You pay 1–2x the time and 2–3x the cost for clearly steadier answers on hard tasks.
Good for
Good for tasks that must be right in one go: code change proposals, structured JSON output, constrained rewrites, and review-style Q&A.
Not for
Not for chat, very short Q&A, or latency-sensitive real-time use. For those, use DeepSeek V4 Flash direct. Use DeepSeek direct →
Why we built it
We can afford the DeepSeek API, but we wanted it a little smarter and more efficient. We are happy to wait twice as long and pay twice as much, as long as the pass rate on hard tasks really goes up. Sansi is the model we wanted for ourselves.
Tuning log
Each row reads: what you expect, whether we delivered, and the evidence. Verdicts use only three words: delivered, partly delivered, not delivered. Numbers come from one reproducible benchmark (7 real editing tickets x 8 runs).
| What you expect | Verdict | Evidence (2026-10-02 benchmark) |
|---|---|---|
| Get it right the first time, fewer redo rounds | Delivered | First-pass success 73% to 89% (41/56 to 50/56), a significant difference (p=0.026) |
| Do not break the JSON or format I asked for | Delivered | When output claims to be JSON but fails to parse, or is empty, the model is asked to resubmit with the error position (at most twice). All 7 to 9 triggers in the benchmark were repaired |
| Do not stop on your own to ask whether to proceed | Delivered | The first version turned execution into a suggestion and asked back in round three. After v2 moved to draft, review, revise-if-defective, this failure is 0 in the benchmark |
| Do not guess when the task contradicts itself | Partly delivered | On self-contradictory tickets, max effort scored 5/8 against 1/8 for the baseline. It is still not 8/8, and we keep adding samples |
| Do not be too slow to use | Partly delivered | Median time is about 1.8x; the longest long task took nearly 15 minutes. For real-time chat use the direct model |
| Do not be absurdly expensive | Delivered | Billed on the tokens you can see, about 3x the direct price. Internal review and revision are not billed separately |
| Do not leave me behind when a new model ships | Delivered (mechanism) | When a new base model ships we run the same benchmark. If direct scores no lower than Sansi, we switch the base and update this page |
| Improve on my real tasks, not only your own tickets | Delivered (mechanism, off by default) | Once you opt in, only failed samples are collected, redacted, kept 90 days, deletable anytime, and fed into the benchmark set every two weeks |
After every tuning round we append a dated line below this table: what changed, which verdicts moved, and the new numbers.
How we tuned it, and the results (measured 2026-10-02, reproducible)
- Tasks: 7 real editing tickets on which a single model loses points (contradictory paths, controlled execution, test completion, repo-level changes). Each configuration ran 8 times, 56 runs in total, scored by objective checkers (exact paths, write scope, syntax check, tests pass), not by eye.
- Baseline: DeepSeek V4 Flash direct at max reasoning effort: 41/56 (73%).
- First version, three questions (fact, inference, next step): 41/56 before the gates, level with the baseline. It only won on tickets whose wording contradicts itself (14/16 vs 9/16), and it added a new failure, an empty body.
- v2, draft, review, revise, plus mechanical gates: 50/56 (89%), p=0.026 against the baseline. Review and revise netted +4 (6 rescued, 2 broken). The gates added +3: when the output is empty or claims to be JSON but fails to parse, the model is asked to resubmit with the error position, at most twice. All 7 to 9 triggers were repaired.
- Cost: median time and spend are about 1.8x.
- Reasoning effort: on ordinary tickets low is as good as max at under half the time and spend; only on self-contradictory tickets is max steadier. So Sansi defaults to high, with four levels available.
Price example
A request with 2,000 input tokens and an 800-token reply costs $0.00108. Internal review and revision rounds are not billed.
Listed prices are final: tax included, no top-up fee, no platform fee.
Use it
Only the base_url and the model id change; the rest is the OpenAI API you already use.
curl https://router.xiaojins.com/v1/chat/completions \
-H "Authorization: Bearer $THRICEGATE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-flash-three-pass",
"reasoning_effort": "high",
"messages": [{"role": "user", "content": "Hello"}]
}'Our commitments
- We test the latest models all the time: when a new model ships we rerun the same benchmark and switch the base if the baseline is stronger.
- Continuous tuning: a MoA (Mixture of Agents) design that balances speed and value. Every change comes with benchmark numbers, published on this page.
- Tune on your tasks: with your consent, failed samples are abstracted into new benchmark tickets, so the next round targets real usage.
Questions
- How is Sansi related to DeepSeek?
- Sansi runs DeepSeek V4.1 Flash three times on your request: it drafts, reviews its own draft as an acceptance tester, and revises only when it finds a defect. Mechanical gates send empty or unparseable output back with the error position.
- Why only user-visible tokens?
- You pay for the tokens you can see: your request and the final reply. The rounds in between are our cost, covered by the higher rate.
- Is it slow?
- About 1.8x DeepSeek direct on the median; the longest benchmark task took close to 15 minutes. For real-time chat use DeepSeek direct.