Two models. One API key each.
Deterministic, self-hosted models with an OpenAI-compatible endpoint. Recharge any amount and start immediately, or take a monthly / yearly plan — either way you pay per token, in ₹, at rates set 20% below the Azure equivalent.
Dhandare 4 Mini
dhandare-4-mini
Chat-completions for structured work: signal narration, post-trade reflection, bull/bear debate, assistant planning with read-only SQL, and answer formatting.
- Endpoint
- https://models.quantforge.co.in/v1
- Input / 10M tokens
- ₹281.60
- Output / 10M tokens
- ₹1126.40
- Cached input / 10M
- ₹70.40
Dhandare Embed
dhandare-embed-1
Sentence embeddings for semantic search and RAG — 384-dimensional vectors, CPU-served.
- Endpoint
- https://embeddings.quantforge.co.in/v1
- Input / 10M tokens
- ₹70.40
- Output / 10M tokens
- ₹0.00
- Cached input / 10M
- ₹0.00
Manthan (Dhandare 3.1 Pro)
dhandare-3.1-pro
Sentence embeddings for semantic search and RAG — 384-dimensional vectors, CPU-served.
- Endpoint
- https://manthan.quantforge.co.in/v1
- Input / 10M tokens
- ₹563.20
- Output / 10M tokens
- ₹1126.40
- Cached input / 10M
- ₹0.00
How billing works
- 1 Recharge any amount and start calling immediately — no plan required. Usage is metered per token and debited in ₹.
- 2 Or take a monthly / yearly plan for a fixed price: it credits the same wallet with more usage credit than it costs, and keeps running for its whole term.
- 3 Every API call checks one thing — that the model has credit left. When it runs out the API says so plainly and top-up takes a moment.
Each model is keyed, priced and credited separately, with its own wallet and its own plans.