Powered by TextCLF Quant

The same models, for less.

Run open and frontier models through one OpenAI-compatible API. Our TextCLF Quant compression cuts the price without touching accuracy — so you ship fast, cheap inference with no GPUs to manage and no lock-in.

$5 free credit for new users

Qwen 3.8 27B

USD / 1M tokens

Input25%
$0.30

others $0.40

Output17%
$2.50

others $3.00

Lower on every token — with the same model quality.

The technology

The same models, for a fraction of the price.

To run big models cheaply, providers quantize them — and most quantize so aggressively that they quietly trade away accuracy. Our TextCLF Quant doesn't: it makes models far cheaper to run while keeping their answers just as good — and we hand every bit of that saving straight back to you as a lower price per token.

4.1x

smaller footprint

99.6%

accuracy retained vs. FP16

TQ 4-bit · Qwen 3.8 27B · WikiText

0.0282

mean KL divergence

92.4%

top-1 agreement

On par with leading quants — but achieved without any calibration data, so the quality carries over to tasks well beyond the benchmark.

  • The most accuracy per bit

    At any given size — say 4 bits per weight — there’s a mathematical best for how much accuracy you can keep. Most methods leave a lot on the table; our TextCLF Quant gets remarkably close to that best, so at the same size our models stay noticeably sharper.

  • Data-free, so it holds up everywhere

    Most methods tune their quant on a calibration dataset, which quietly overfits to whatever it was tuned on — great on that test, shakier everywhere else. Ours uses no calibration data at all, so the quality holds steady across the full range of tasks, not just the ones it was measured on.

  • Same answers, proven

    We test every model against the untouched original and keep it within a fraction of a percent on standard benchmarks — no silent quality drops.

  • The savings go straight to you

    Smaller models are cheaper for us to run — and we don’t pocket the difference. Every bit of efficiency TextCLF Quant unlocks is passed directly to you as a lower price per token.

One API. Every model worth running.

Switch models with a single string. Every model on TextCLF is served from the same OpenAI-compatible endpoint at the same low, transparent per-token price.

  • Llama 3.3 70B Instruct

    128K context

    Meta

    Input $0.08
    Output $0.30
    Cached input
    6.6 tok/s
  • Llama 3.1 8B Instruct

    128K context

    Meta

    Input $0.018
    Output $0.038
    Cached input
    21.2 tok/s
  • Qwen3.8 27B

    262K context

    Alibaba

    Input $0.30
    Output $2.50
    Cached input $0.03
    21.3 tok/s
  • Qwen3 Coder Next

    262K context

    Alibaba

    Input $0.10
    Output $0.78
    Cached input $0.03
  • DeepSeek V4 FlashSoon

    1M context

    DeepSeek

    Input $0.05
    Output $0.12
    Cached input $0.01
  • MiMo V2.5Soon

    context

    Xiaomi MiMo

    Input $0.11
    Output $0.23
    Cached input $0.00255

Prices shown per 1M tokens (USD). 100+ additional models available.

View full model catalog →

Change one line. Cut your bill.

Already using the OpenAI SDK? Point the base URL at TextCLF and keep the rest of your code exactly as it is.

  • Drop-in replacement for the OpenAI SDK
  • Same request and response schema
  • Streaming, tools, and JSON mode supported
quickstart.sh
curl https://api.textclf.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_TEXTCLF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3.8-27B",
    "messages": [
      { "role": "user", "content": "Explain inference cost in one line." }
    ]
  }'

Prepaid credit. Zero surprises.

Get $5 in free credit to start — no credit card needed. When you're ready for more, top up your balance and spend it at each model's per-token rate. You can never be charged more than you've added.

Free credit

$5on us

Start building the moment you sign up — no credit card required.

  • $5 in credit, free
  • No credit card to start
  • Access to every model
  • OpenAI-compatible API
Claim free credit

Prepaid credits

Most popular
Top upfrom $10

Add credit and draw it down as you call models. Each model bills at its own per-token rate — you only spend what you use.

  • Per-token pricing, set per model
  • Credits never expire
  • Auto-reload so you never run dry
  • Real-time usage & spend dashboard
Buy credits

Enterprise

Customvolume

Committed-use rates, dedicated capacity, and invoicing.

  • Volume discounts on credit
  • Dedicated GPU capacity
  • SOC 2 Type II, SSO/SAML
  • Invoicing & dedicated support
Contact sales

Every model has its own input and output rate. See per-token pricing for all models.