LLM7 model guide
Streaming models
Chat models that report incremental response streaming.
Requirements and limitations
Handle partial chunks, interrupted connections, and final usage separately. Streaming does not imply lower total latency.
Only active chat models with explicit support are included. Missing capability values remain unknown; top-level values take precedence over nested values. Image input also establishes vision support.
| Model | Pricing | Context | Tool calling |
|---|---|---|---|
| $0.02 USD input and $0.04 USD output per 1M tokens | 400,000 | Supported | |
| $1.00 USD input and $4.05 USD output per 1M tokens | 512,000 | Not supported | |
| $0.50 USD input and $1.20 USD output per 1M tokens | 512,000 | Supported | |
| $0.04 USD input and $0.05 USD output per 1M tokens | 8,000 | Not supported | |
| $0.40 USD input and $2.00 USD output per 1M tokens | 256,000 | Supported | |
| $1.00 USD input and $3.00 USD output per 1M tokens | 1,024,000 | Supported | |
| $4.50 USD input and $30.00 USD output per 1M tokens | 1,000,000 | Supported | |
| $4.00 USD input and $15.00 USD output per 1M tokens | 1,000,000 | Supported | |
| $0.04 USD input and $0.20 USD output per 1M tokens | 200,000 | Supported | |
| $2.25 USD input and $12.19 USD output per 1M tokens | 1,000,000 | Supported | |
| $2.50 USD input and $12.50 USD output per 1M tokens | 1,000,000 | Supported | |
| $0.12 USD input and $0.45 USD output per 1M tokens | 1,000,000 | Supported | |
| $0.45 USD input and $2.25 USD output per 1M tokens Minimum: $0.000784/request | 1,000,000 | Supported | |
| $0.01 USD input and $0.02 USD output per 1M tokens | 32,000 | Supported | |
| $0.06 USD input and $0.18 USD output per 1M tokens | 1,024,000 | Supported | |
| $0.18 USD input and $0.60 USD output per 1M tokens | 1,000,000 | Supported | |
| $1.11 USD input and $3.33 USD output per 1M tokens | 1,000,000 | Supported | |
| $0.03 USD input and $0.08 USD output per 1M tokens | 1,048,576 | Supported | |
| $0.02 USD input and $0.04 USD output per 1M tokens | 256,000 | Supported | |
| $0.06 USD input and $0.30 USD output per 1M tokens | 1,000,000 | Supported | |
| $0.05 USD input and $0.15 USD output per 1M tokens | 1,000,000 | Supported | |
| gemma4:31b | $0.07 USD input and $0.23 USD output per 1M tokens | 262,000 | Supported |
| $0.50 USD input and $2.00 USD output per 1M tokens | 1,000,000 | Supported | |
| $0.15 USD input and $0.55 USD output per 1M tokens | 1,048,576 | Supported | |
| $0.20 USD input and $1.00 USD output per 1M tokens Minimum: $0.0001/request | 1,050,000 | Supported | |
| $0.45 USD input and $0.58 USD output per 1M tokens | Unknown context | Supported | |
| $1.92 USD input and $5.40 USD output per 1M tokens | 1,000,000 | Supported | |
| $0.24 USD input and $1.00 USD output per 1M tokens | 1,000,000 | Supported | |
| $5.00 USD input and $10.00 USD output per 1M tokens | 1,000,000 | Supported | |
| $0.30 USD input and $1.00 USD output per 1M tokens | 500,000 | Supported | |
| $0.40 USD input and $0.50 USD output per 1M tokens | 500,000 | Supported | |
| $2.00 USD input and $10.00 USD output per 1M tokens | 1,000,000 | Supported | |
| llama-4-maverick | $0.19 USD input and $0.75 USD output per 1M tokens | 1,048,576 | Supported |
| $0.03 USD input and $0.03 USD output per 1M tokens | 128,000 | Not supported | |
| $0.06 USD input and $0.08 USD output per 1M tokens | 32,000 | Not supported | |
| $0.10 USD input and $0.40 USD output per 1M tokens | 250,000 | Supported |
Example request
POST https://api.llm7.io/v1/chat/completions
Authorization: Bearer YOUR_LLM7_API_KEY
Content-Type: application/json
{
"model": "DeepSeek-V4-Flash-0731",
"messages": [
{
"role": "user",
"content": "Hello!"
}
],
"stream": true
}Explore related pages
Model catalogModel comparisonsCapabilitiesIntegrationsCost calculatorsToken calculatorTool calling modelsVision modelsLong context modelsJSON mode modelsReasoning modelsHermes Agent with LLM7OpenClaw with LLM7n8n with LLM7LangChain with LLM7Vercel AI SDK with LLM7Customer support cost calculatorDocument processing cost calculator