Models & pricing
Open-weight models served from 8× NVIDIA V100 32 GB in India. Prices are per 1,000,000 tokens in INR — no forex, no card surcharges, GST invoice included.
| Model | Input ₹/1M | Output ₹/1M | Context | Live speed | Status |
|---|---|---|---|---|---|
|
Llama 3.3 70B
llama-3.3-70b
Flagship general chat, tensor-parallel over NVLink |
65.00 | 78.00 | 131,072 | 17.4 tok/s | ready |
|
Qwen3 Coder 30B
qwen3-coder-30b
Code generation and agentic coding. MoE, 3B active - fast. |
16.00 | 64.00 | 65,536 | 13 tok/s | ready |
|
Qwen2.5-VL 7B Instruct
qwen2.5-vl-7b-instruct
Vision-language: images in, text out. Also handles plain text chat. |
21.00 | 21.00 | 98,304 | 51.3 tok/s | ready |
|
Llama 3.1 8B
llama-3.1-8b
Fast volume tier |
11.00 | 11.00 | 98,304 | 22.5 tok/s | ready |
|
Qwen3 8B
qwen3-8b
Fast tier with reasoning |
4.50 | 15.00 | 40,960 | 31.2 tok/s | ready |
|
BGE-M3
bge-m3
embedding Multilingual embeddings, 1024-dim |
1.10 | — | 8,192 | — | ready |
|
MiniLM L6 v2
all-minilm-l6-v2
embedding Lightweight embeddings, 384-dim |
1.50 | — | 256 | — | ready |
Cost calculator
Per day
—
Per month
—
Start with free trial credit
OpenAI-compatible at https://api.bookmyhost.com/v1.
Change two lines of code and you're on Indian infrastructure, billed in rupees.
