Pick an engine
Each preset is a page: what it runs, what each part costs, and the questions it answers. Or compose your own from the same price book.
Last updated
The default pipeline for English. Deepgram Nova-3 transcribes, GPT-4o-mini via OpenRouter reasons, and ElevenLabs speaks: Turbo v2.5 for English, Flash v2.5 for other languages. It is the cheapest composed engine and the one to start with.
Indian languages end to end. Sarvam Saaras v3 transcribes, the Sarvam-105B model reasons, and Bulbul v3 speaks, so Hindi, Tamil, Telugu and the other Indian languages Sarvam supports are handled by models built for them rather than by an English pipeline with a language flag.
One model for the whole turn. Google Gemini Live takes the caller's audio in and produces audio out with the reasoning inside the provider, so there is no separate speech-to-text or text-to-speech stage to name and latency is what the model itself delivers.
Deepgram's Voice Agent API runs the conversation as a single realtime session and drives ElevenLabs Turbo v2.5 for the voice itself. The synthesised text is not exposed, so its character count is inferred from the transcript and carries a small safety margin in the price.
Name a speech-to-text, language and text-to-speech model from one family and see the minute they add up to.
The full price list: per second, per token and per character, for every model a composed engine can name.


