9 INTENT MODELS · MEMORY · YOUR OPENROUTER KEY

Route cheap when cheap.
Remember the rest.

OpenAI Chat Completions, OpenAI Responses, and Anthropic /v1/messages all classify each turn into a lane, trim waste context, and keep premium models in reserve. You pay OpenRouter for tokens. Strix sells the router and the memory layer.

~80%estimated model savings vs always-Sonnet
Representative 5,000-request agent workload with 27.2% less payload, compared with Claude Sonnet 4.6 list pricing. Fast lanes (quick, coding) carry everyday turns. Your results vary.
OpenAI Chat + Responses Anthropic /v1/messages You bring OpenRouter
Works with
  • Cursor
  • Claude Code
  • Codex
  • OpenCode
  • Cline
  • Roo Code
  • Continue
  • Aider
01 / WHAT YOU BUY

Save. Remember.
Stay in control.

Not a model catalog and not a token reseller. Strix is the decision layer between your coding clients and the providers you already pay.

01

Save money

Fast and coding lanes take cheap turns. Deep waits for hard work. Payload filtering drops stale history before it is billed.

02

Remember

Facts are extracted and injected silently across sessions and editors, scoped per project. Pin, recall, history, feedback, batch, and MCP on the same key. Toggle, export, or forget from the dashboard.

03

Control

Builder or Team. Keys, seats after three, and spend visibility. Your OpenRouter key stays yours; Team is the SKU when more than one person routes.

02 / THE SYSTEM

Nine intent models.
One stable contract.

Point Cursor, Claude Code, Codex, OpenCode, Cline, Roo Code, Continue, or Aider at one base URL. Strix classifies, filters, and egresses.

01 / REQUESTAgent intent arrivesMessages, tools, mode, and task evidence
02 / CONTROLContext is right-sizedLive evidence stays; dead weight can leave
03 / ROUTECapability meets costHealth-aware providers with escalation ready
01

Intent routing

Classifies work into nine public intent models — coding, quick, deep, bug-hunt, visual, writing, research, uncensored — or leave the client on auto.

02

Payload efficiency

Removes stale context and oversized schemas while preserving the evidence your model actually needs.

03

Resilient egress

Health-aware provider ordering, circuit state, streaming pass-through, and automatic transient failover.

04

Bring your key

OpenRouter BYOK is required. We do not mark up those tokens. Optional Venice only if you store a Venice key.

03 / SPEND LESS

Your coding tool sends
more than you think.

Agents resend history, tool schemas, and large tool results on every turn. A premium model then prices every retained token. Strix cuts both sides: send less, then route remaining work to the lowest-cost capable lane instead of always-Sonnet.

01

Send fewer tokens

Drops stale history, duplicate context, oversized tool results, and idle schemas before they become billable input.

02

Buy only what the step needs

Quick and coding stay efficient. Advanced intelligence is reserved for difficult reasoning and escalation.

03

Prove the difference

See tokens kept, tokens avoided, baseline cost, optimized cost, and savings by day, client, mode, and capability tier.

LIVE ESTIMATE

See the coding-agent bill before and after Strix.

Change request volume and workload shape. The baseline uses the same Sonnet 4.6 list price used by the product evidence pipeline; optimized routes use prices already tracked in this repository.

Workload
ESTIMATED MONTHLY CODING-AGENT SAVINGS$33382% less model spend

At 5,000 requests per month, this scenario drops the modeled bill from $405 to $72.

You keep$333estimated monthly savings
Baseline$405Sonnet 4.6 · $3 input / $15 output per MTok
Optimized$72filtered input + scenario route
Input avoided4,896tokens per request

27.2% measured Agent payload reduction in Strix telemetry. Uses 18,000 input and 1,800 output tokens per request. Estimate, not a guarantee; provider prices and workload mix change.

Pricing source: Anthropic Claude Platform pricing, accessed July 2026. Filter band source: repository evidence constants and measured Agent telemetry.

04 / WHAT CHANGES

The goal stays.
The dead weight does not.

Strix does not summarize away the task. It preserves the current goal, relevant evidence, required tool protocol, and forced tools. It reduces categories of context that do not help this step.

USER PROMPT
Refactor provider failover, preserve streaming, and add regression tests.
ROUTE → refactor-safe
PRESERVED
  • + latest user goal
  • + relevant file context
  • + recent tool evidence
  • + edit tools
ELIGIBLE TO REDUCE
  • idle browser/MCP schemas
  • duplicate messages
  • oversized old tool output
  • agent boilerplate

Required tools and matched tool-call history remain attached.

05 / BEHIND THE ROUTER

Real models.
One stable interface.

Strix maintains a changing, health-checked range of capable model families. Your client keeps one configuration while routing adapts to the work, availability, and cost.

01
AnthropicFrontier reasoning
02
OpenAIFast + frontier coding
03
GoogleMultimodal work
04
DeepSeekAgentic coding
05
Z.aiArchitecture + review
06
MiniMaxSpecialist recovery
WHY A RANGE?

Fast utility work should not pay for Sol or Opus. Auto stays on cheap capable lanes. Frontier models stay in reserve unless you opt in.

ONE PUBLIC CONTRACT

You choose task-shaped lanes. Strix keeps provider identities, fallback order, and parameter compatibility behind the routing layer.

06 / GET STARTED

Start a 7-day trial.
Bring your OpenRouter key.

Create a workspace, copy your Strix key once, store OpenRouter (tokens stay on their bill), then point any supported client at the same base URL.

7-day trial · Builder $29/mo · Team $99/mo · OpenRouter BYOK
  1. 01
    Create your account

    Builder for one person, Team when you need seats after three.

  2. 02
    Store OpenRouter, copy Strix

    Tokens bill to OpenRouter. The full Strix key appears once — copy it then.

  3. 03
    Connect any client

    Cursor, Claude Code, Codex, OpenCode, Cline, Roo Code, Continue, or Aider. Same OpenAI-compatible contract.

07 / PRICING

Builder $29. Team $99.
You bring OpenRouter.

7-day trial on both plans. No markup on BYOK tokens. Strix charges the subscription plus request and memory-derive overage.

7-DAY TRIALBUILDER

One person, every public lane, your OpenRouter key

$29 / month

7-day trial. You bring OpenRouter — they bill tokens; Strix does not mark up BYOK inference.

  • 7-day trial · OpenRouter key required
  • 1 user · 2 API keys · memory (5 projects)
  • 50,000 routed requests / month included
  • 2,000 memory derives / month included
  • Overage $0.40 / 1k requests · $0.002 / derive
  • Nine public models · OpenAI + Anthropic APIs
Start 7-day trial
7-DAY TRIALTEAM

Shared workspace with extra seats after the first three

$99 / month

Three seats included, then $20 per extra seat. Same BYOK model: you pay OpenRouter for tokens.

  • 7-day trial · 3 seats · $20 / extra seat
  • 250,000 routed requests / month included
  • 2,000 memory derives / month included
  • Overage $0.40 / 1k requests · $0.002 / derive
  • Shared keys, spend controls, and memory (50 projects)
  • Optional managed inference: 15% on OpenRouter cost if you have no stored key
Start Team trial
COST COMPARISON

See which plan fits.

Enter expected native OpenRouter usage. Builder is $29/mo; Team is $99/mo. Managed inference adds 15% only if Strix runs on the platform key (no stored OpenRouter credential).

LOWER ESTIMATED CUSTOMER COSTBYOKBased only on funding cost; both modes route identically.
BYOK total$129.00
Managed total$144.00
Native provider cost$100.00
Managed Strix fee (15%)$15.00

BYOK totals are OpenRouter cost plus the Strix subscription. Managed inference adds 15% on native OpenRouter cost only when there is no stored key. Request and derive overage is metered separately.

FAQ

Straight answers
before you start.

Trial needs an OpenRouter key. Tokens stay on their bill. Strix is the router.

Do I bring my own OpenRouter key?
Yes. The 7-day trial and paid plans require an OpenRouter key (BYOK). OpenRouter bills model tokens at their rates. Strix sells routing, memory, and the control plane — not a token markup on BYOK inference.
What am I paying Strix for?
Nine public intent models, payload filtering, failover, a memory layer (cross-session facts, Cohere rerank, forget/export), and Team seats/keys/spend controls. Auto plus cheap lanes such as quick and coding keep everyday turns off always-Sonnet pricing.
Which clients work?
Any OpenAI-compatible client, plus Anthropic /v1/messages. That includes Cursor, Claude Code, Codex, OpenCode, Cline, Roo Code, Continue, and Aider — not Cursor only. Point the base URL at api.strixgate.dev and use a strix_* key.
Is uncensored or Venice required?
No. Uncensored is an optional lane. Venice is used only if you store a Venice key. Default routing stays on OpenRouter with your key.
Do you log prompts?
No prompt logging for training or human review of your request bodies. We meter usage, lanes, and spend. You can forget or export stored memory facts from the dashboard.
How is this different from OpenRouter or Cursor alone?
OpenRouter still bills tokens — you bring that key. Strix sits in front with lanes, filtering, and memory. Cursor is one supported client, not the only one. Use Compare in the nav for OpenRouter, Cursor, and alternatives pages.
How does Team pricing work?
Team is $99/month with three seats included, then $20 per extra seat. Builder is $29/month for one person. Both include a 7-day trial.