Intent routing
Classifies work into specialized lanes for coding, deep reasoning, review, research, writing, and visual engineering.
Preparing the control plane…
Put Cursor—or any OpenAI-compatible coding tool—behind one endpoint. Strix removes billable context waste and avoids paying Claude Sonnet-class prices for every step, while preserving frontier capability when the work needs it.
route_00491Correct capability. Lower-cost model.
Not another model catalog. Strix is the decision layer between your tools and every provider you trust.
Classifies work into specialized lanes for coding, deep reasoning, review, research, writing, and visual engineering.
Removes stale context and oversized schemas while preserving the evidence your model actually needs.
Health-aware provider ordering, circuit state, streaming pass-through, and automatic transient failover.
Compares the optimized request with a consistent premium baseline, then separates filtering, routing, and cache savings in your dashboard.
Cursor and other agents resend history, tool schemas, and large tool results on every turn. Then a premium model can price every retained input and output token. Strix cuts both sides of that bill: send less, then route the remaining work to the lowest-cost capable tier.
Drops stale history, duplicate context, oversized tool results, and idle schemas before they become billable input.
Fast utility work stays on an efficient tier. Advanced intelligence is reserved for difficult reasoning and escalation.
See tokens kept, tokens avoided, baseline cost, optimized cost, and savings by day, client, mode, and capability tier.
Change request volume and workload shape. The baseline uses the same Sonnet 4.6 list price used by the product evidence pipeline; optimized routes use prices already tracked in this repository.
At 5,000 requests per month, this scenario drops the modeled bill from $405 to $72.
27.2% measured Agent payload reduction in Strix telemetry. Uses 18,000 input and 1,800 output tokens per request. Estimate, not a guarantee; provider prices and workload mix change.
Pricing source: Anthropic Claude Platform pricing, accessed July 2026. Filter band source: repository evidence constants and measured Agent telemetry.
Strix does not summarize away the task. It preserves the current goal, relevant evidence, required tool protocol, and forced tools. It reduces categories of context that do not help this step.
“Refactor provider failover, preserve streaming, and add regression tests.”ROUTE → refactor-safe
Required tools and matched tool-call history remain attached.
Strix routes across a changing, health-checked range of capable models. The exact model can change by task, availability, price, and your chosen lane—without changing your Cursor configuration.
Fast utility work should not pay for frontier reasoning. Difficult tasks can escalate, and unhealthy providers can fail over. Public lanes describe capability, not a promise that one upstream model will be used forever.
Strix does not add a separate topic-moderation layer or rewrite what you are allowed to ask. Infrastructure protections still block secret theft, prompt extraction, abuse, and cost attacks; upstream provider policies continue to apply.
Start with a free Strix account. We prepare your workspace, issue your first API key once, and then walk you through connecting Cursor or another compatible coding client.
No card required · setup takes about three minutesSet up the workspace that owns your usage, keys, and settings.
The full key appears after setup. Copy it then—it cannot be recovered.
Open Cursor Models directly, then add Auto and Quick as your everyday and fast modes.
Every plan includes routing, filtering, measured savings, and provider access managed by Strix at a transparent plan rate.
Enter your expected underlying model spend and seats. We calculate the subscription, inference fee, seat charges, and total.
Strix manages provider access and bills measured inference cost plus the plan rate. Taxes are not included.