Models & pricing

Open-weight models served from 8× NVIDIA V100 32 GB in India. Prices are per 1,000,000 tokens in INR — no forex, no card surcharges, GST invoice included.

Model Input ₹/1MOutput ₹/1M Context Live speed Status
Llama 3.3 70B
llama-3.3-70b
Flagship general chat, tensor-parallel over NVLink
65.00 78.00 131,072 17.4 tok/s ready
Qwen3 Coder 30B
qwen3-coder-30b
Code generation and agentic coding. MoE, 3B active - fast.
16.00 64.00 65,536 13 tok/s ready
Qwen2.5-VL 7B Instruct
qwen2.5-vl-7b-instruct
Vision-language: images in, text out. Also handles plain text chat.
21.00 21.00 98,304 51.3 tok/s ready
Llama 3.1 8B
llama-3.1-8b
Fast volume tier
11.00 11.00 98,304 22.5 tok/s ready
Qwen3 8B
qwen3-8b
Fast tier with reasoning
4.50 15.00 40,960 31.2 tok/s ready
BGE-M3
bge-m3 embedding
Multilingual embeddings, 1024-dim
1.10 8,192 ready
MiniLM L6 v2
all-minilm-l6-v2 embedding
Lightweight embeddings, 384-dim
1.50 256 ready

Cost calculator

Per day
Per month
Start with free trial credit

OpenAI-compatible at https://api.bookmyhost.com/v1. Change two lines of code and you're on Indian infrastructure, billed in rupees.