Skip to content
VaakyoDocs
Navigation
Open console →

Agents

Guardrails

The Vaakyo platform team sets guardrails: rules every agent on the platform must follow, in every workspace. They are added to each agent's prompt, and every phone call is checked after it ends. If an agent breaks a rule, the workspace gets a strict notice. If the same agent breaks a rule again within the window, the agent is frozen until the platform team unfreezes it.

The rules

These are the default rules for India. The platform team can change them, so GET /api/v1/guardrails/rules always returns the current list. The Rules button on the console’s Compliance page opens this page.

RuleWhat is not allowedSeverity
Impersonating government officials or policeClaiming to be (or to call for) a government department, the police, a court, RBI, TRAI, CBI, customs or income tax when the business isn’t that authority.violation
Asking for OTPs, passwords, card or bank detailsAsking for an OTP, password, PIN, CVV, full card number, net-banking login or UPI PIN, or asking the caller to install a screen-sharing app.violation
Threats, abuse or harassmentThreatening, insulting or harassing the caller, or using abusive, sexist, casteist or communal language.violation
Coercion or false urgencyUsing threats of arrest, legal action, account blocking or fake deadlines to make the caller pay or act.violation
Political campaigning without disclosureAsking for votes or making election promises without saying which party or candidate the call is for.violation
Adult contentSexual or sexually suggestive content.violation
Misleading financial or medical claimsGuaranteed returns, loan approval promised whatever the caller’s eligibility, cures, diagnoses or prescriptions.violation
Ignoring a request to stopThe caller asks not to be called again, or to stop, and the agent keeps pitching.violation
Claiming to be humanThe caller sincerely asks whether they’re talking to a person, and the agent says it’s human.warning
Calling outside 9 am to 9 pm (TRAI)An outbound call started outside 9:00–21:00 IST. This is a clock check, not an LLM judgement. TRAI allows service and transactional calls at any time, so this rule only warns.warning

Only what the agent said counts. If the caller is abusive, that isn’t a violation by your agent. Refusing a request, quoting a rule, or warning the caller about fraud doesn’t count either.

How calls are checked

  • During the call. The enabled conversation rules are added to the end of every agent’s system prompt under # Platform rules. They override anything above them in the prompt.
  • After the call. When a phone call ends, the platform’s model reads the transcript at the same time as the agent’s post-call analytics. For each rule broken, it returns a short exact quote of the agent’s words, an explanation and a confidence score. Only violations at or above the platform’s threshold count (80% by default). If the quote isn’t actually in the transcript, its confidence is halved.
  • Browser test calls and agent tests aren’t checked.
  • Cost. The check runs on a small model and costs a fraction of a paisa per call. The platform pays for it; it never appears in a call’s cost_breakdown.

Each violation becomes an incident. The call record gets guardrail_incidents (how many incidents) and guardrail_severity (violation or warning). The call’s timeline also gets a guardrail.violation event.

Strikes, notices and freezing

  • Only rules with violation severity count as strikes. Warnings are listed on the Compliance page and sent as events, but they never count toward freezing.

  • Strikes are counted per agent, one per call: two rules broken on the same call are one strike. Only strikes from the last 30 days count, and dismissed incidents never count. The platform team can change both settings.

  • First strike: a strict notice. A red banner sits at the top of every console page until someone acknowledges the notice. The workspace’s owners and admins also get an email with the agent, the rule, the quote and a link to the call.

  • Second strike: the agent is frozen. Its status becomes "frozen", and frozen_reason, frozen_at and frozen_incident_id are set. A frozen agent:

    • can’t place calls. POST /api/v1/agents/{id}/call returns 409, and any of its calls still in the queue are canceled.
    • doesn’t answer its number. Callers hear a short “not available” message, then the call ends.
    • has its running and scheduled campaigns paused, and they can’t be resumed or created.
    • can’t be duplicated. Saving it keeps it frozen, and you can’t set status: "frozen" yourself.
    • still works for browser test calls, so you can check a fixed prompt.

    The owners, admins and the platform team are emailed.

  • Only the platform team can unfreeze an agent. This includes Composer and API keys: they can’t. Unfreezing puts the agent back to its previous status. Its paused campaigns stay paused until you resume them.

Acknowledge or appeal

On the Compliance page, a person signed in to the console (owner, admin or developer) can:

  • Acknowledge a notice. This clears the banner; the strike still counts.
  • Appeal an incident with a note explaining why the check got it wrong, or what you changed. The platform team reviews the appeal. They can dismiss it (it stops being a strike, and you’re emailed), confirm it, or unfreeze the agent.

API keys, connected apps and Composer can read incidents but can’t acknowledge or appeal them.

EndpointWhat it does
GET /api/v1/guardrails/rulesThe enabled rules, window_days and strikes_to_freeze
GET /api/v1/guardrails/summaryOpen notices (open, latest) and frozen_agents (what the banner shows)
GET /api/v1/guardrails/incidentsYour incidents, newest first. Filters: status, agent_id, call_id, severity, limit, skip
GET /api/v1/guardrails/incidents/{id}One incident
POST /api/v1/guardrails/incidents/{id}/acknowledgeAcknowledge (signed-in people only)
POST /api/v1/guardrails/incidents/{id}/appealAppeal, with {"note": "..."} (10–2,000 characters; signed-in people only)

An incident looks like this:

{
  "id": "gi_3f9c1a7e2b4d4c0a9e1f",
  "agent_id": "a1b2c3",
  "agent_name": "Loan reminder",
  "call_id": "c9d8e7",
  "rule_id": "sensitive_credentials",
  "rule_name": "Asking for OTPs, passwords, card or bank details",
  "severity": "violation",
  "quote": "Please tell me the OTP you just received.",
  "explanation": "The agent asked the caller for an OTP.",
  "confidence": 0.95,
  "status": "open",
  "action": "notice",
  "strike": 1,
  "appeal": null,
  "resolution": null,
  "created_at": "2026-10-04T15:42:10+05:30"
}

status is open, acknowledged, appealed, confirmed or dismissed. action is notice (a strike), frozen (the strike that froze the agent) or warning.

Webhook

Subscribe an endpoint to guardrail.violation to hear about incidents as they happen (see Webhook events). Each incident sends one event, with data: incident_id, rule_id, rule_name, severity, quote, confidence, action and strike.

Esc