valency-one · hybrid inference

Your agent thinks in the cloud. Your GPU sits idle.

valency-one distributes inference between the cloud and your local compute. The cloud model does the reasoning. Your machine writes the answer.

up to 4.4× fewer cloud tokens. keep near-full accuracy.

valency-one sport

100.0%of cloud accuracy, on average*

valency-one eco

91.4%of cloud accuracy, on average*

* Averaged over seven benchmarks: MATH500, AMC23, AIME 2024, AIME 2025, GPQA-Diamond, HumanEval+ and MBPP+, each against Qwen3-4B with thinking on.

01Cloud–local

The cloud reasons. Your machine answers.

The cloud model does the reasoning your local models can’t. Writing the answer is offloaded to your device, cutting cloud inference cost while keeping accuracy close to the full model.

Joint reinforcement learning for hybrid inference
CLOUDreasons
Qwen3-4B4B · cloud
  • Qwen3-4B4B · cloudactive
  • Mimo V2.6 Flashcloudcoming next
sends only what’s needed
YOUR MACHINEwrites the answer
Qwen3-0.6B0.6B · local
  • Qwen3-0.6B0.6B · localactive
  • Qwen 3.8 27B27B · localcoming next

100.0% of cloud accuracy−49% cloud cost

02Use your device’s compute

Full model vs. hybrid inference

One real question from MATH500. Both answer it correctly. One of them bills you for every token of its thinking.

You

Qwen3-4B alone, with thinking

−√3 < x < √3

3,554 billed cloud tokens
valency-one

−√3 < x < √3

732billed cloud tokens454on your machine, free

−79% cloud cost for the same correct answer.

03Benchmarks

Same accuracy. A fraction of the cloud costs.

Accuracy

0%25%50%75%100%MATHSCIENCECODINGMATH500: 89.3%MATH500: 90.4%MATH500: 87.2%MATH500AMC23: 88.8%AMC23: 87.5%AMC23: 88.8%AMC23AIME 2024: 63.3%AIME 2024: 63.3%AIME 2024: 48.3%AIME24AIME 2025: 51.7%AIME 2025: 54.2%AIME 2025: 40.8%AIME25GPQA-Diamond: 51.4%GPQA-Diamond: 51.3%GPQA-Diamond: 49.4%GPQAHumanEval+: 89.0%HumanEval+: 87.1%HumanEval+: 84.5%HE+MBPP+: 78.3%MBPP+: 76.6%MBPP+: 75.3%MBPP+
Sport matches the full model; eco trades a little accuracy for the deepest savings.

Cloud cost savings

Math
MATH5004,821 tokens
AMC237,571 tokens
AIME 202411,756 tokens
AIME 202512,964 tokens
Science
GPQA-Diamond6,285 tokens
Coding
HumanEval+3,275 tokens
MBPP+2,536 tokens
32–77% fewer cloud tokens than the full model, depending on mode and task.

Full model = Qwen3-4B with thinking on. 4 samples per question.

04Questions

What you might be wondering

Is valency-one available today?

It is in early access. The results on this page come from our first trained pair: Qwen3-4B in the cloud and Qwen3-0.6B on your machine. Join the list and we’ll email you when your spot opens.

What runs on my machine?

A small local model that reads the cloud’s reasoning and writes the final answer. Its tokens run on hardware you already own, so they never appear on a cloud bill.

Does it change the answers I get?

Sport mode matches the full model’s accuracy on average across seven math, science and coding benchmarks. Eco mode spends the least cloud and keeps 91% of the accuracy on average.

How is this different from a smaller model?

The two models are trained together with reinforcement learning, so the cloud learns to send just enough reasoning for your machine to finish the answer correctly. A small model alone scores far lower.

Early access

Use the compute you already have.

Be the first to experience hybrid inference.

Also from mortiphi

Mortic

Talk to your OpenCode coding agent. Ask codebase questions and hear answers in your current session. Developer preview for Apple Silicon Macs.