Agents
Guardrails
The Vaakyo platform team sets guardrails: rules every agent on the platform must follow, in every workspace. They are added to each agent's prompt, and every phone call is checked after it ends. If an agent breaks a rule, the workspace gets a strict notice. If the same agent breaks a rule again within the window, the agent is frozen until the platform team unfreezes it.
The rules
These are the default rules for India. The platform team can change them, so GET /api/v1/guardrails/rules always returns the current list. The Rules button on the console’s Compliance page opens this page.
| Rule | What is not allowed | Severity |
|---|---|---|
| Impersonating government officials or police | Claiming to be (or to call for) a government department, the police, a court, RBI, TRAI, CBI, customs or income tax when the business isn’t that authority. | violation |
| Asking for OTPs, passwords, card or bank details | Asking for an OTP, password, PIN, CVV, full card number, net-banking login or UPI PIN, or asking the caller to install a screen-sharing app. | violation |
| Threats, abuse or harassment | Threatening, insulting or harassing the caller, or using abusive, sexist, casteist or communal language. | violation |
| Coercion or false urgency | Using threats of arrest, legal action, account blocking or fake deadlines to make the caller pay or act. | violation |
| Political campaigning without disclosure | Asking for votes or making election promises without saying which party or candidate the call is for. | violation |
| Adult content | Sexual or sexually suggestive content. | violation |
| Misleading financial or medical claims | Guaranteed returns, loan approval promised whatever the caller’s eligibility, cures, diagnoses or prescriptions. | violation |
| Ignoring a request to stop | The caller asks not to be called again, or to stop, and the agent keeps pitching. | violation |
| Claiming to be human | The caller sincerely asks whether they’re talking to a person, and the agent says it’s human. | warning |
| Calling outside 9 am to 9 pm (TRAI) | An outbound call started outside 9:00–21:00 IST. This is a clock check, not an LLM judgement. TRAI allows service and transactional calls at any time, so this rule only warns. | warning |
Only what the agent said counts. If the caller is abusive, that isn’t a violation by your agent. Refusing a request, quoting a rule, or warning the caller about fraud doesn’t count either.
How calls are checked
- During the call. The enabled conversation rules are added to the end of every agent’s system prompt under
# Platform rules. They override anything above them in the prompt. - After the call. When a phone call ends, the platform’s model reads the transcript at the same time as the agent’s post-call analytics. For each rule broken, it returns a short exact quote of the agent’s words, an explanation and a confidence score. Only violations at or above the platform’s threshold count (80% by default). If the quote isn’t actually in the transcript, its confidence is halved.
- Browser test calls and agent tests aren’t checked.
- Cost. The check runs on a small model and costs a fraction of a paisa per call. The platform pays for it; it never appears in a call’s
cost_breakdown.
Each violation becomes an incident. The call record gets guardrail_incidents (how many incidents) and guardrail_severity (violation or warning). The call’s timeline also gets a guardrail.violation event.
Strikes, notices and freezing
-
Only rules with
violationseverity count as strikes. Warnings are listed on the Compliance page and sent as events, but they never count toward freezing. -
Strikes are counted per agent, one per call: two rules broken on the same call are one strike. Only strikes from the last 30 days count, and dismissed incidents never count. The platform team can change both settings.
-
First strike: a strict notice. A red banner sits at the top of every console page until someone acknowledges the notice. The workspace’s owners and admins also get an email with the agent, the rule, the quote and a link to the call.
-
Second strike: the agent is frozen. Its
statusbecomes"frozen", andfrozen_reason,frozen_atandfrozen_incident_idare set. A frozen agent:- can’t place calls.
POST /api/v1/agents/{id}/callreturns409, and any of its calls still in the queue are canceled. - doesn’t answer its number. Callers hear a short “not available” message, then the call ends.
- has its running and scheduled campaigns paused, and they can’t be resumed or created.
- can’t be duplicated. Saving it keeps it frozen, and you can’t set
status: "frozen"yourself. - still works for browser test calls, so you can check a fixed prompt.
The owners, admins and the platform team are emailed.
- can’t place calls.
-
Only the platform team can unfreeze an agent. This includes Composer and API keys: they can’t. Unfreezing puts the agent back to its previous status. Its paused campaigns stay paused until you resume them.
Acknowledge or appeal
On the Compliance page, a person signed in to the console (owner, admin or developer) can:
- Acknowledge a notice. This clears the banner; the strike still counts.
- Appeal an incident with a note explaining why the check got it wrong, or what you changed. The platform team reviews the appeal. They can dismiss it (it stops being a strike, and you’re emailed), confirm it, or unfreeze the agent.
API keys, connected apps and Composer can read incidents but can’t acknowledge or appeal them.
| Endpoint | What it does |
|---|---|
GET /api/v1/guardrails/rules | The enabled rules, window_days and strikes_to_freeze |
GET /api/v1/guardrails/summary | Open notices (open, latest) and frozen_agents (what the banner shows) |
GET /api/v1/guardrails/incidents | Your incidents, newest first. Filters: status, agent_id, call_id, severity, limit, skip |
GET /api/v1/guardrails/incidents/{id} | One incident |
POST /api/v1/guardrails/incidents/{id}/acknowledge | Acknowledge (signed-in people only) |
POST /api/v1/guardrails/incidents/{id}/appeal | Appeal, with {"note": "..."} (10–2,000 characters; signed-in people only) |
An incident looks like this:
{
"id": "gi_3f9c1a7e2b4d4c0a9e1f",
"agent_id": "a1b2c3",
"agent_name": "Loan reminder",
"call_id": "c9d8e7",
"rule_id": "sensitive_credentials",
"rule_name": "Asking for OTPs, passwords, card or bank details",
"severity": "violation",
"quote": "Please tell me the OTP you just received.",
"explanation": "The agent asked the caller for an OTP.",
"confidence": 0.95,
"status": "open",
"action": "notice",
"strike": 1,
"appeal": null,
"resolution": null,
"created_at": "2026-10-04T15:42:10+05:30"
}
status is open, acknowledged, appealed, confirmed or dismissed. action is notice (a strike), frozen (the strike that froze the agent) or warning.
Webhook
Subscribe an endpoint to guardrail.violation to hear about incidents as they happen (see Webhook events). Each incident sends one event, with data: incident_id, rule_id, rule_name, severity, quote, confidence, action and strike.