> ## Documentation Index
> Fetch the complete documentation index at: https://docs.openference.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Model Request Ratios

> What each model costs against your plan allowance, and how many requests that buys on every plan.

# Model Request Ratios

Not every model costs the same slice of your plan. Each model has a **ratio** —
how much of your allowance one call to it uses.

**Auto and flash models are the 1.0 floor.** GLM-5.2 costs **2.0** — one
call uses two requests' worth. Your plan's "requests per window" number is what
that buys on Auto and other 1.0 models; GLM-5.2 calls consume twice
as much. On **Free** (**25** per 5-hour window) that is **25** Auto calls or
**12** GLM-5.2 calls; a very long GLM-5.2 prompt costs **4.0**
(down to **6** on Free).

> **Note:** mid-tier models sit between 1.0 and 2.0; premium chat models cost
> more — see the table below. GLM-5.2 long-context surcharges are
> separate — see [Long-context surcharge](#long-context-surcharge). Image models
> use a fixed ratio per image — see [Image generation ratios](#image-generation-ratios).
> **Promo ratio:** `Qwen3.8 27b` is **0.5** through **2026-09-30** UTC (standard ratio ~~1~~).
> **Promo ratio:** `DeepSeek-V4-Flash-0731` is **0.75** through **2026-09-30** UTC (standard ratio ~~1~~).

## Ratios at a glance

| Ratio          | Models                                                      |
| -------------- | ----------------------------------------------------------- |
| **4**          | `GLM-5.2`, `GLM-5.3` when the prompt exceeds 262,144 tokens |
| **2**          | `GLM-5.2`, `GLM-5.3`                                        |
| **1.7**        | `DeepSeek-V4-Pro-0813`                                      |
| **1.6**        | `Kimi K2.7 Code`                                            |
| **1**          | `Auto`, `GLM-5.3-Flash`, `MiniMax M3`                       |
| ~~1~~ **0.75** | `DeepSeek-V4-Flash-0731`                                    |

Use the model name exactly as shown as the `model` value in your request.

Tables below list models on subscription plans (Free through Max+). On-demand-only
models (for example Kimi K3) are billed per token and are not shown here.

> **Legacy Auto Agent plans:** Existing subscribers on Auto Agent Mini, Lite,
> Pro, or Max keep their grandfathered allowance at the same ratios until they
> choose to switch plans.

## How many requests you get

### Per 5-hour window

| Model                          | Ratio          | Free          | Lite            | Pro               | Pro+                | Max                 | Max+                |
| ------------------------------ | -------------- | ------------- | --------------- | ----------------- | ------------------- | ------------------- | ------------------- |
| **GLM-5.2** (prompt > 262,144) | **4**          | 6             | 100             | 200               | 300                 | 400                 | 800                 |
| **GLM-5.3** (prompt > 262,144) | **4**          | 6             | 100             | 200               | 300                 | 400                 | 800                 |
| GLM-5.2                        | 2              | 12            | 200             | 400               | 600                 | 800                 | 1,600               |
| GLM-5.3                        | 2              | 12            | 200             | 400               | 600                 | 800                 | 1,600               |
| DeepSeek-V4-Pro-0813           | 1.7            | 14            | 235             | 470               | 705                 | 941                 | 1,882               |
| Kimi K2.7 Code                 | 1.6            | 15            | 250             | 500               | 750                 | 1,000               | 2,000               |
| Auto                           | 1              | 25            | 400             | 800               | 1,200               | 1,600               | 3,200               |
| GLM-5.3-Flash                  | 1              | 25            | 400             | 800               | 1,200               | 1,600               | 3,200               |
| MiniMax M3                     | 1              | 25            | 400             | 800               | 1,200               | 1,600               | 3,200               |
| DeepSeek-V4-Flash-0731         | ~~1~~ **0.75** | ~~25~~ **33** | ~~400~~ **533** | ~~800~~ **1,066** | ~~1,200~~ **1,600** | ~~1,600~~ **2,133** | ~~3,200~~ **4,266** |

### Per week

Weekly caps reset Monday 00:00 UTC.

| Model                          | Ratio          | Free            | Lite                 | Pro                   | Pro+                  | Max                   | Max+                  |
| ------------------------------ | -------------- | --------------- | -------------------- | --------------------- | --------------------- | --------------------- | --------------------- |
| **GLM-5.2** (prompt > 262,144) | **4**          | 87              | 1,875                | 3,250                 | 4,875                 | 6,500                 | 13,000                |
| **GLM-5.3** (prompt > 262,144) | **4**          | 87              | 1,875                | 3,250                 | 4,875                 | 6,500                 | 13,000                |
| GLM-5.2                        | 2              | 175             | 3,750                | 6,500                 | 9,750                 | 13,000                | 26,000                |
| GLM-5.3                        | 2              | 175             | 3,750                | 6,500                 | 9,750                 | 13,000                | 26,000                |
| DeepSeek-V4-Pro-0813           | 1.7            | 205             | 4,411                | 7,647                 | 11,470                | 15,294                | 30,588                |
| Kimi K2.7 Code                 | 1.6            | 218             | 4,687                | 8,125                 | 12,187                | 16,250                | 32,500                |
| Auto                           | 1              | 350             | 7,500                | 13,000                | 19,500                | 26,000                | 52,000                |
| GLM-5.3-Flash                  | 1              | 350             | 7,500                | 13,000                | 19,500                | 26,000                | 52,000                |
| MiniMax M3                     | 1              | 350             | 7,500                | 13,000                | 19,500                | 26,000                | 52,000                |
| DeepSeek-V4-Flash-0731         | ~~1~~ **0.75** | ~~350~~ **466** | ~~7,500~~ **10,000** | ~~13,000~~ **17,333** | ~~19,500~~ **26,000** | ~~26,000~~ **34,666** | ~~52,000~~ **69,333** |

## Image generation ratios

Each successful image counts as one call. Ratios apply per image, not per token.

| Model          | Ratio |
| -------------- | ----- |
| SDXL Base 1.0  | 25    |
| SDXL Lightning | 25    |

You do not have to spend a window on one model — costs simply add up. On Pro,
600 GLM-5.2 calls (1,200) plus 400 Auto calls (400) is 1,600 requests' worth.

## Long-context surcharge

**A GLM-5.2 request whose prompt exceeds 262,144 tokens costs 4.0 instead of 2.0.**

* GLM-5.2 is the **only** chat model with a context surcharge.
* It is measured on **input** only. A long *reply* to a short prompt still costs 2.0.
* The threshold is measured on **upstream-reported input tokens** after the
  call completes. `X-Quota-Cost` on the response uses a pre-flight estimate and
  may read lower until settlement; the dashboard is always exact.

If you routinely work above 262k tokens, `DeepSeek-V4-Pro-0813` (alias `DeepSeek-V4-Pro`) and `MiniMax M3`
both have larger context windows and carry **no** surcharge.

## Reading the cost at runtime

You never have to hardcode this table.

`GET /v1/models` reports the ratio on every model:

```bash theme={null}
curl https://api.openference.com/v1/models \
  -H "Authorization: Bearer YOUR_API_KEY"
```

```json theme={null}
{
  "id": "GLM-5.2",
  "quota_multiplier": 2,
  "quota_context_surcharge": {
    "threshold_input_tokens": 262144,
    "factor": 2
  }
}
```

`quota_context_surcharge` is present only on models whose cost scales with
prompt size.

Every response also carries what that specific call used:

| Header              | Meaning                                                    |
| ------------------- | ---------------------------------------------------------- |
| `X-Quota-Cost`      | What this request cost, in requests (e.g. `1`, `2`, `2.5`) |
| `X-Quota-Remaining` | Approximately what is left in the current window           |

> **Note:** `X-Quota-Remaining` is read when the request is admitted, so with
> several requests in flight at once it can read slightly high. The dashboard
> under **Billing** is always exact.

## What still counts as one request

The ratio changes how much a call costs, not which calls count:

* `success` and `client_error` (4xx) count.
* **Upstream provider errors and capacity-unavailable situations do not count** —
  a failed request is refunded in full, at whatever it would have cost.
* The **per-minute burst limit** is unaffected. It is a hard cap on calls per
  minute and is never extended by credits.

## Related

* [Plans & Usage](/billing/plans) — allowances, weekly caps, burst limits
* [Pricing & Transparency](/models/pricing) — on-demand token rates beyond your allowance
* [Model Catalog](/models/catalog) — context windows and capabilities
