中文

Models

Sansi: a MoA model that makes DeepSeek a little smarter

The same DeepSeek V4 Flash: draft first, review, and revise only if a defect is found, with mechanical gates on top. You pay 1–2x the time and 2–3x the cost for clearly steadier answers on hard tasks.

Why we built it

We can afford the DeepSeek API, but we wanted it a little smarter and more efficient. We are happy to wait twice as long and pay twice as much, as long as the pass rate on hard tasks really goes up. Sansi is the model we wanted for ourselves.

How we tuned it, and the results (measured 2026-10-02, reproducible)

  • Tasks: 7 real editing tickets on which a single model loses points (contradictory paths, controlled execution, test completion, repo-level changes). Each configuration ran 8 times, 56 runs in total, scored by objective checkers (exact paths, write scope, syntax check, tests pass), not by eye.
  • Baseline: DeepSeek V4 Flash direct at max reasoning effort: 41/56 (73%).
  • First version, three questions (fact, inference, next step): 41/56 before the gates, level with the baseline. It only won on tickets whose wording contradicts itself (14/16 vs 9/16), and it added a new failure, an empty body.
  • v2, draft, review, revise, plus mechanical gates: 50/56 (89%), p=0.026 against the baseline. Review and revise netted +4 (6 rescued, 2 broken). The gates added +3: when the output is empty or claims to be JSON but fails to parse, the model is asked to resubmit with the error position, at most twice. All 7 to 9 triggers were repaired.
  • Cost: median time and spend are about 1.8x.
  • Reasoning effort: on ordinary tickets low is as good as max at under half the time and spend; only on self-contradictory tickets is max steadier. So Sansi defaults to high, with four levels available.

Tuning log

Each row reads: what you expect, whether we delivered, and the evidence. Verdicts use only three words: delivered, partly delivered, not delivered. Numbers come from one reproducible benchmark (7 real editing tickets x 8 runs).

What you expectVerdictEvidence (2026-10-02 benchmark)
Get it right the first time, fewer redo roundsDeliveredFirst-pass success 73% to 89% (41/56 to 50/56), a significant difference (p=0.026)
Do not break the JSON or format I asked forDeliveredWhen output claims to be JSON but fails to parse, or is empty, the model is asked to resubmit with the error position (at most twice). All 7 to 9 triggers in the benchmark were repaired
Do not stop on your own to ask whether to proceedDeliveredThe first version turned execution into a suggestion and asked back in round three. After v2 moved to draft, review, revise-if-defective, this failure is 0 in the benchmark
Do not guess when the task contradicts itselfPartly deliveredOn self-contradictory tickets, max effort scored 5/8 against 1/8 for the baseline. It is still not 8/8, and we keep adding samples
Do not be too slow to usePartly deliveredMedian time is about 1.8x; the longest long task took nearly 15 minutes. For real-time chat use the direct model
Do not be absurdly expensiveDeliveredBilled on the tokens you can see, about 3x the direct price. Internal review and revision are not billed separately
Do not leave me behind when a new model shipsDelivered (mechanism)When a new base model ships we run the same benchmark. If direct scores no lower than Sansi, we switch the base and update this page
Improve on my real tasks, not only your own ticketsDelivered (mechanism, off by default)Once you opt in, only failed samples are collected, redacted, kept 90 days, deletable anytime, and fed into the benchmark set every two weeks

After every tuning round we append a dated line below this table: what changed, which verdicts moved, and the new numbers.

Where it fits, and where it does not

Good for tasks that must be right in one go: code change proposals, structured JSON output, constrained rewrites, and review-style Q&A.

Not for chat, very short Q&A, or latency-sensitive real-time use. For those, use DeepSeek V4 Flash direct.

Our commitments

  • We test the latest models all the time: when a new model ships we rerun the same benchmark and switch the base if the baseline is stronger.
  • Continuous tuning: a MoA (Mixture of Agents) design that balances speed and value. Every change comes with benchmark numbers, published on this page.
  • Tune on your tasks: with your consent, failed samples are abstracted into new benchmark tickets, so the next round targets real usage.

How to use it

OpenAI-compatible. Set model to deepseek-flash-three-pass. reasoning_effort accepts low, medium, high, or max, and max_tokens defaults to 65,536. Response headers carry per-round time and tokens, and the console shows three cards (draft, review, revise).

curl https://router.xiaojins.com/v1/chat/completions \
  -H "Authorization: Bearer $XJP_KEY" -H "content-type: application/json" \
  -d '{
    "model": "deepseek-flash-three-pass",
    "reasoning_effort": "high",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Pricing details and API key requestsQuick start