valency-one · hybrid inference
Your agent thinks in the cloud. Your GPU sits idle.
valency-one distributes inference between the cloud and your local compute. The cloud model does the reasoning. Your machine writes the answer.
up to 4.4× fewer cloud tokens. keep near-full accuracy.
100.0%of cloud accuracy, on average*
91.4%of cloud accuracy, on average*
* Averaged over seven benchmarks: MATH500, AMC23, AIME 2024, AIME 2025, GPQA-Diamond, HumanEval+ and MBPP+, each against Qwen3-4B with thinking on.
01Cloud–local
The cloud reasons. Your machine answers.
The cloud model does the reasoning your local models can’t. Writing the answer is offloaded to your device, cutting cloud inference cost while keeping accuracy close to the full model.
Qwen3-4B4B · cloud▾
- Qwen3-4B4B · cloudactive
- Mimo V2.6 Flashcloudcoming next
Qwen3-0.6B0.6B · local▾
- Qwen3-0.6B0.6B · localactive
- Qwen 3.8 27B27B · localcoming next
02Use your device’s compute
Full model vs. hybrid inference
One real question from MATH500. Both answer it correctly. One of them bills you for every token of its thinking.
xxx
−√3 < x < √3
−√3 < x < √3
−79% cloud cost for the same correct answer.
03Benchmarks
Same accuracy. A fraction of the cloud costs.
Accuracy
Cloud cost savings
Full model = Qwen3-4B with thinking on. 4 samples per question.
04Questions
What you might be wondering
Is valency-one available today?
It is in early access. The results on this page come from our first trained pair: Qwen3-4B in the cloud and Qwen3-0.6B on your machine. Join the list and we’ll email you when your spot opens.
What runs on my machine?
A small local model that reads the cloud’s reasoning and writes the final answer. Its tokens run on hardware you already own, so they never appear on a cloud bill.
Does it change the answers I get?
Sport mode matches the full model’s accuracy on average across seven math, science and coding benchmarks. Eco mode spends the least cloud and keeps 91% of the accuracy on average.
How is this different from a smaller model?
The two models are trained together with reinforcement learning, so the cloud learns to send just enough reasoning for your machine to finish the answer correctly. A small model alone scores far lower.
Early access
Use the compute you already have.
Be the first to experience hybrid inference.
Also from mortiphi
Mortic
Talk to your OpenCode coding agent. Ask codebase questions and hear answers in your current session. Developer preview for Apple Silicon Macs.