Agents
An agent is the AI persona that runs each conversation. Configure the script, language, voice, models, call behavior, analytics, and inbound phone assignment from one builder.
Create an agent
- Open Agent Setup from the sidebar.
- Click New agent. You are asked How do you want to start? — pick Auto Build to describe the call and have it written for you, or Start from Template to begin from a ready-made agent (or the blank Custom template).
- Rename the agent from the summary header so it is easy to select later in campaigns and batches.
- Fill the Agent, LLM, Audio, Engine, Call, Analytics, Inbound, and Tools tabs.
- Click Save from the header actions before using the agent at scale.
Build with AI (Auto Build)
Rather than writing everything by hand, choose Auto Build in the New agent chooser and describe the screening call in plain English. The system generates a complete agent — welcome message, full call prompt, and a matching rating template — that you review, tweak, and save.
- Give the agent a title and pick its spoken name (e.g. Isha) — the name the agent says on the call.
- In Describe the screening call, type what you need (role, skills, number of questions, what to evaluate). Not sure what to write? Click Examples (top-right) to drop in a ready-made description.
- Optionally fill the structured fields: Calling from (the company the agent names in the opener — defaults to Veytrix AI), Role, Client / hiring company, Number of questions, and Focus areas. All optional — they sharpen the result.
- Click Generate, review the welcome / prompt / rating template, then Save this agent.
- Finish by attaching a phone number and a voice.
Agent tab
This tab controls the conversation purpose.
- Agent type - choose Real conversation for two-way AI calls, or Announcement for one-way calls that play the welcome message and hang up.
- Welcome message - the first line spoken to the caller. Write it as fixed text; voice agents don't substitute CSV columns, so don't use
{name}-style placeholders (they are read aloud literally). - Agent prompt - the core instruction set: goal, tone, required questions, disqualification rules, and closing behavior.
You are a screening agent for a Java Full Stack role.
Greet the candidate by name.
Confirm availability for the role and location.
Ask about Java, Spring Boot, React, SQL, and project experience.
Ask one follow-up when an answer is unclear.
Keep each response under two sentences.
If the candidate asks a policy question, answer only from the linked knowledge base.
End politely after all required questions are complete.The conversation pipeline
Veytrix streams call audio through three pluggable providers, each chosen per agent:
| Stage | Providers | Role |
|---|---|---|
| STT (speech-to-text) | Sarvam, Deepgram, Cartesia | Transcribes the candidate in real time. |
| LLM | Gemini, Groq, OpenAI, Cerebras | Decides what the agent says next (sentence-streamed). |
| TTS (text-to-speech) | Cartesia, Sarvam, ElevenLabs | Speaks the agent's reply. |
LLM tab
- Provider and model - choose OpenAI, Groq, Cerebras, or Gemini, then select a supported model.
- API credential - select the workspace credential from Settings > AI Services. Hidden on managed workspaces, which use our keys — see step 2 of Getting started.
- Tokens - maximum generated tokens for each response. Higher values allow longer replies but can slow the call.
- Temperature - lower values are more consistent; higher values are more varied.
- Knowledge bases - multi-select reference documents that the agent can use during the call.
Audio tab
- Language - the language the agent speaks. Eleven are available: English, Hindi, Bengali, Tamil, Telugu, Kannada, Malayalam, Marathi, Gujarati, Punjabi, and Odia.
- Also understand and speak - extra languages the caller and agent may use. Sarvam detects each utterance and the reply is spoken in that language.
- Speech-to-text - choose Sarvam, Deepgram, or Cartesia and select the transcription model.
- Keywords - add important names, skills, products, or acronyms to improve recognition.
- Text-to-speech - choose Cartesia, Sarvam, or ElevenLabs, then pick a voice from the searchable voice picker, plus model, speed, and credential.
- Voice quality - for Sarvam, choose telephony, standard, or high-quality sample rate based on the model.
Conversation engines
- Pipeline (default) — the modular STT → LLM → TTS chain above. Maximum control; ~0.8–2s response.
- Gemini-Live — Google's speech-to-speech model; lowest latency, fewer knobs.
Engine tab
- Pipeline uses separate STT, LLM, and TTS settings. Use it when you need provider control, knowledge bases, and detailed tuning.
- Gemini Live uses one speech-to-speech model. Use it when low latency is more important than separate provider control.
- Interrupt after words controls how quickly the agent can interrupt after the caller starts speaking.
- Endpointing controls how quickly the system treats caller speech as complete.
- Response rate lets you trade speed against smoother delivery.
Call tab
Use this tab to control telephony and call safety behavior.
| Field | What it does |
|---|---|
| Telephony provider | Currently Plivo. |
| Account / Number | Uses the default Plivo account or a selected credential/number. |
| Outbound number | Which of the workspace’s own numbers this agent dials from — the caller ID the candidate sees. Leave it empty to follow the workspace default. Useful when one workspace runs several agents that should each call from a different line. |
| Ambient noise | Adds optional office/cafe/street-style background noise. |
| Noise cancellation | Attempts to reduce caller-side background noise. |
| Voicemail detection | Helps identify voicemail instead of a live candidate. |
| Record calls | Records only when the platform recording switch and agent toggle are both enabled. |
| Auto Reschedule | Automatically re-dials no-answer / failed calls one more time. |
| Auto Callback | When a candidate asks to be called back at a specific time (e.g. “tomorrow around 12” or “after 20 minutes”), the agent automatically dials them again at that time after the call is rated. Off by default. |
| Final call message | What the agent says before ending the call. |
| Hangup on user silence | Ends the call when the caller stays silent longer than the configured seconds. |
| Total call timeout | Maximum total call length. |
Analytics tab
- Call rating template decides which rubric and columns are used on the Ratings page.
- Analytics webhook URL sends execution data to your own endpoint.
- Summarization creates a post-call summary.
- Extraction captures custom structured fields from the transcript.
Inbound tab
Assign an inbound-capable number to the agent. One inbound number can ring exactly one agent, and numbers already assigned to other agents are hidden. Manage the full number list from My Numbers.
Tools tab — Google Calendar
Turn on Google Calendar to auto-create a calendar event (with a Google Meet link) shortly after a call is rated — using the date and time the candidate confirmed during the call. The event is created on a host’s Workspace calendar and the candidate is invited.
| Field | What it does |
|---|---|
| Calendar account | Which connected Google Calendar account to use (or the org default). Add accounts under Integrations. |
| Host email | The Workspace user whose calendar the event is created on (e.g. recruiter@yourdomain). Must be in your Google Workspace. |
| Event name | Event title template — supports personalization like {name} or {role}. |
| Event duration | Length of the created event, in minutes. |
Key settings
| Setting | Default | Effect |
|---|---|---|
| Hang up on silence | 20s | Ends the call after this much candidate silence. |
| Call timeout | 600s | Hard cap on total call length. |
| LLM temperature | 0.2 | Lower = more consistent answers. |
| Recording | off | Per-agent opt-in (also needs the platform recording switch on). |
| Summarization / Extraction | off | Post-call summary and structured field extraction. |
| Rating template | screening | The rubric used to score the call. |
Try the agent before you dial anyone
Three ways to hear or read an agent, all from the header actions in Agent Setup:
| Action | What happens | Costs |
|---|---|---|
| Web Call (Beta) | Talk to the agent in your browser using your microphone — no phone, no number needed. It runs the same pipeline a real call runs: speech-to-text, the LLM, the voice, and barge-in all behave exactly as they will on the phone. | Nothing |
| Get a call from agent | The agent dials a real phone number you enter, so you hear it over the actual telephony leg. Pick the country beside the number field, then enter the number. | Wallet, at the normal per-minute rate |
| Chat with agent | Types the conversation instead of speaking it. The quickest way to check what the prompt makes the agent say. | Nothing |
Test and publish
- Save the agent.
- Start a Web Call for a fast, free pass over the whole conversation.
- Then use Get a call from agent with your own phone number, to hear it over a real line.
- Check greeting, pronunciation, language, interruption timing, and final message.
- Open Call History to inspect transcript and summary.
- Open Ratings to confirm the selected rating template produces the expected columns.
- After validation, select this agent when creating a campaign or batch.