API v1 · DEVELOPER BETA

Build calling into your application

Use your backend to control calls, invite people, receive events, and request optional recording analysis.

Base URL: https://api.kallavoice.com

Live carrier access is restricted to approved test accounts. A sandbox key does not activate a SIP connection, enable arbitrary destinations, or incur a subscription charge.

Move an application from Twilio

  1. Verify your Kalla account, choose features, and create a backend API key. GSE's complimentary pilot waives Kalla billing only; it does not activate a carrier or waive outside provider charges.
  2. In Application settings, select the key, your JSON event callback URL, and your incoming-call TwiML URL. Save the signing secret once in your backend secret manager. Authorize any additional callback or audio-service HTTPS origins in the same form. WSS audio endpoints use their corresponding HTTPS origin for registration.
  3. Add your website origin under Browser calling. Download the server SDK, callback verifier, call-status reducer, conference-status reducer, browser audio SDK, signed-in session adapter, and cross-tab ownership adapter. The account page exports your non-secret application configuration.
  4. Connect a separate test SIP trunk/number, submit its routing limits, and validate it. Kalla verifies carrier ownership and SIP/media interoperability before activation. Do not move your production number during setup.
  5. Test inbound, outbound, queue, voicemail, transfers, group calls, mobile and failure recovery in parallel with your existing provider. Keep a provider switch in your application. Change your production carrier routing only after acceptance, and retain the old settings for rollback.

What changes in your application

Twilio integrationKalla integration
Account SID/auth tokenAccount-scoped Bearer API key. Base URL: https://api.kallavoice.com
Calls, conferences, participants, recordingsTwilio-style server SDK resources. Asynchronous controls return receipts; distinguish accepted work from completed native effects.
Voice TwiML responseSupported Voice TwiML behaviors listed below. Validate your real scripts. Unsupported options are rejected rather than ignored.
X-Twilio-SignatureKalla V2 callback verifier. Verify raw bytes before parsing; persist event IDs to deduplicate retries and reject stale state updates.
Twilio Device/access tokenKallaAudio and restricted call grants issued by your backend. This is a browser adapter change, not a token-name substitution.
Call and recording callbacksPreserve your business rules behind a provider adapter. An ended leg does not by itself prove the customer answered.

Continue after a voice engine disconnects

With <Connect action="/after-engine" method="POST"><Stream url="wss://voice.example.com/audio"/></Connect>, Kalla drains and removes the stream, then requests the registered action URL with signed call parameters. Its TwiML response replaces the remaining instructions. Without action, the existing script continues after the stream closes. Put cloud-owned routing or voicemail there for engine outages. A cancelled/replaced script or ended call does not invoke the old action. This integrates transport only; your voice engine remains separate.

Recover after a callback outage

The account page has a Callback recovery section for failed or blocked script status callbacks. Your key needs Call history and Routing configuration to inspect delivery, plus Call controls to retry. The retry rechecks the current destination and permissions, preserves the original event ID and payload, and returns pending rather than pretending delivery succeeded. Scrubbed or revoked transcription payloads cannot be revived. Your receiver must deduplicate event IDs; a prior request may have reached it even when Kalla did not receive its response. JSON event callbacks retry separately through their persistent cursor.

Backend equivalents: GET /v1/notifications?status=failed and POST /v1/notifications/{id}/retry with an empty JSON object and Idempotency-Key. Lists contain at most 100 records and omit event payloads. Use the same idempotency key if retrying a lost recovery response.

Conference audio and recording callbacks

Conference jitterBufferSize accepts small (20 ms), medium (40 ms), large (60 ms, default), or off. It applies to each participant independently and is removed when that participant leaves. These sizes contribute to total call latency. The region attribute selects a configured media region. Kalla currently offers de1 (Germany); other requested regions return conference_region_unavailable. Conference eventCallbackUrl is supported for older integrations: a signed POST with StatusCallbackEvent=conference-record-end is sent after the room ends and its recording finishes. Prefer recordingStatusCallback; supplying it suppresses the deprecated callback.

Conference callback ordering

Bind verified callbacks to the exact tenant, conference ID, room name, and callback URL stored by your application. The conference reducer rejects callbacks from an older room generation. Persist an event inbox before acknowledging delivery; retain sequence gaps until missing events arrive. Commit the reducer state and your business projection atomically, with idempotent outbox commands for external effects. Do not invoke a nontransactional webhook inside a database transaction. Persistent gaps require reconciliation, not silently skipping participant transitions.

Signed-in browser lifecycle

KallaVoiceSession wraps KallaAudio with single-flight connection, selected-call grant checks, and cleanup when the signed-in user changes or the application loses call ownership. Supply your authenticated identity getter, immediate identity-change subscription, and atomic cross-tab ownership adapter. Identity strings do not grant permissions: your backend still authorizes every grant. Ownership acquisition returns a synchronous release function for that exact lease; notify onLost if another tab takes over. Close or dispose the session on sign-out. Closing releases this browser leg, not the parent customer call. Download and bundle the browser modules with your application. The optional createBrowserOwnership adapter uses Web Locks to arbitrate same-origin tabs atomically. Give it an application/tenant scope and bridge your existing call coordinator through isBlocked, publish, and subscribe. It fails closed without Web Locks; use a backend lease adapter on unsupported browsers. Backend claims are still required across devices.

Audio devices

After microphone permission, list available devices with the browser mediaDevices API. KallaAudio and KallaVoiceSession provide selectInputDevice(deviceId) and selectOutputDevice(deviceId). Input changes keep mute and noise suppression, stop the old microphone after successful replacement, and do not redial. A denied change retains the working input. Output selection uses the browser permission API and reports unsupported browsers explicitly; mobile output routing remains browser-controlled.

Browser audio

Your backend authenticates the CRM user, checks their role and the call they may join, then requests POST /v1/calls/{callId}/grants with their subject ID. Return only that call's token to the browser. Never expose the tenant API key. Register the application's exact HTTPS origin in the Kalla account.

import { KallaAudio } from './kalla-browser.mjs';
const audio = new KallaAudio(document.querySelector('audio'), console.log);
// Run from an explicit user gesture; your backend supplies callId/token.
await audio.connect({
  endpoint: 'https://media.kallavoice.com/rtc/offer',
  callId, token
});
// audio.mute(true), await audio.sendDigit('1'), audio.close()

The user grants microphone permission. Provide an explicit playback button if Safari blocks autoplay. On logout or navigation, close audio and revoke the application's call authority. Browser-origin removal also closes its active audio legs. Full iPhone/carrier acceptance is a separate check from desktop browser fixtures.

1. Create a scoped API key

Sign in to your verified account, enable the features you need, and open Application API keys. Give the key a name, select its permissions and expiry, and save the secret when it appears. Kalla shows the secret only once. Create a replacement key before revoking an old one when rotating credentials.

Keep the key on your backend. Do not put it in browser JavaScript, mobile bundles, URLs, analytics or logs. General API requests with an Origin header are rejected. Browser audio uses short-lived grants issued by your backend instead.

curl https://api.kallavoice.com/v1/capabilities \
  -H "Authorization: Bearer $KALLA_API_KEY"

Capabilities require both the key's permission and the account's enabled feature. Revocation or removing a feature also blocks queued operations before dispatch. Keys expire after 1, 7, 30 or 90 days; your account can hold up to ten unexpired active keys.

GET /v1/capabilities returns tenant, keyId, enabled, available and apiVersion. Account sign-in and key management use your session cookie and CSRF protection; an API key cannot manage another account.

2. Create and control a call

Mutations use JSON. Call creation, commands, programs, invitations and analysis jobs require an Idempotency-Key. Reuse that value only when retrying the same request. Changed content returns a conflict. A 202 response means accepted, not answered.

curl https://api.kallavoice.com/v1/calls \
  -H "Authorization: Bearer $KALLA_API_KEY" \
  -H 'Content-Type: application/json' \
  -H 'Idempotency-Key: example-call-001' \
  -d '{"to":"+12025550101","from":"+12025550102","browserSubject":"staff-123"}'

The numbers above are reserved examples and work only in an operator-configured fixture tenant. Use only your account's approved destinations and caller IDs. Permissions: calls.create; browser preparation also needs browser.grants. Optional context is {"type":"customer"|"user","id":"your-record-id"}. Your backend authenticates the user and authorizes access to that record.

The receipt contains id, callId and status. Poll GET /v1/commands/{id} until completed. An uncertain result may already have performed the action: do not create another call to retry it.

EndpointRequired permission / behavior
GET /v1/calls
GET /v1/calls/{call}
calls.read. Call state, version, participant mapping and retained context.
POST /v1/calls/{call}/commandscalls.control plus operation-specific permission. Body: operation, version, args.
GET /v1/commands/{command}calls.read. Completed, rejected and uncertain outcomes remain distinguishable.
POST /v1/calls/{call}/grantsbrowser.grants. Body: {"subject":"authenticated-user-id"}. Returns one-use token, callId, subject and expiry.
GET /v1/routing
PUT /v1/routing
Also configurable under Call queues in the account console.
routing. PUT: version, queues, aiAdapters, fallback. Current fallback value: voicemail.

Commands: hangup, hold, resume, connect_destination use empty args. add_participant takes destination; remove_participant takes participantId (both require conferences). queue takes queueId and optional timeoutSeconds; pickup_queue takes destination and subject (queues). voicemail takes empty args (voicemail). The live driver does not yet accept ai_handoff.

{"operation":"hangup","version":3,"args":{}}

Read the current call version before a new command. On a version conflict, inspect state before deciding what to do next; never blindly redial.

Start recording an active call

Send POST /v1/calls/{callId}/recordings with an Idempotency-Key and optional recordingTrack (inbound, outbound, or both), recordingChannels (mono or dual), trim, and recordingStatusCallback settings. The server SDK exposes client.calls(callId).recordings.create(options). Poll GET /v1/recording-starts/{id}; the returned recordingId identifies its metadata, controls, and authenticated audio. Kalla attaches a receive-only tap to the caller's native audio channel without replacing the call script or moving the caller. The tap follows that channel into a conference, stops when the call ends or the initiating key loses access, and sends signed recording callbacks. One live recording per call is supported. Set playBeep:true to play a beep to the recorded call before capture begins; a failed beep does not silently start recording. The default is false. Dual recordings put incoming caller audio on the left and outgoing audio on the right; single-track selections remain mono.

Pause, resume, or stop recording

Send POST /v1/recordings/{recordingId} with {"status":"paused"}, {"status":"in-progress"}, or {"status":"stopped"}. Use an Idempotency-Key and poll the returned control at GET /v1/recording-controls/{controlId}. A completed pause means native capture is paused; do not begin a private conversation on an accepted/pending receipt alone. These controls require calls.control, voicemail, and recordings.read; conference recordings additionally require conferences. An uncertain pause triggers a stop attempt and is never replayed automatically. Stopping a Record script lets it continue to its next step. The same controls apply to both stereo channels; pauseBehavior selects silence or skip.

Leave Dial with the star key

Set hangupOnStar="true" on Dial to let its caller press * after admission and continue with the Dial action or next script step. The default is false. The called party's keypresses do not control this escape. Number, SIP, browser, mixed, queue, and conference paths use their ordinary cleanup so recordings, invited media, and shared-room membership drain before continuation.

Conference roster

List current members and pending native invitations with GET /v1/conferences/{id}/participants. A leg awaiting answer or press-one admission appears as connecting; joined means it has entered the room. Fetch one participant by call ID or exact label at GET /v1/conferences/{id}/participants/{reference}, or client.conferences(id).participants(reference).fetch(). Ambiguous references are rejected. Pending participants must count when deciding whether a room is abandoned. Updates, removal, and private announcements also accept the exact label. The accepted operation stores the resolved call ID, so retrying its idempotency key never retargets a replacement participant using that label.

Change conference exit behavior

Update an active participant with boolean endConferenceOnExit and/or beepOnExit. The change applies when that participant subsequently leaves; it does not move or interrupt their current audio. Poll the returned conference-control receipt before relying on the new setting. Subscribe to modify events for signed changes. Ending a conference releases its participants to their own Dial continuation.

Conference announcements

Send POST /v1/conferences/{conferenceId}/announcements with announceUrl and optional announceMethod, or use client.conferences(id).update({announceUrl, announceMethod}). The URL must be on an authorized callback origin and return bounded Say/Play/Pause/Redirect TwiML or supported audio. Each announcement has a two-minute total limit. Poll its receipt at GET /v1/conference-announcements/{id}; accepted work is not completed playback. Subscribe to announcement conference events for signed announcement-end/announcement-fail callbacks. Participants keep their existing scripts and calls. For a private announcement, use POST /v1/conferences/{conferenceId}/participants/{callId}/announcements or client.conferences(id).participants(callId).update({announceUrl, announceMethod}). A temporary whisper-only audio tap sends it to that participant without moving them out of the room. The private completion callback includes that CallSid when still present. A new announcement replaces the previous one for the same audience after its playback drains. Private announcements to different participants remain independent. The superseded receipt is marked cancelled; completed playback retains its own receipt.

Recording privacy pauses

Update a live recording with {"status":"paused","pauseBehavior":"silence"} to preserve the timeline with silent audio, or pauseBehavior:"skip" to omit that interval. Silence is the default. Resume with {"status":"in-progress"}. Mode changes while paused establish the new privacy guard before releasing the old one. Wait for the control receipt to complete before collecting sensitive information. Uncertain capture controls stop recording; an unconfirmed stop ends the owned call. No uncertain request is automatically replayed.

Conference speaker activity

Include speaker in the first participant's conference statusCallbackEvent subscription to receive signed participant-speech-start and participant-speech-stop callbacks. They identify the participant and conference, share the conference sequence number, and contain no transcript or audio. Held, muted, and private coaching speech is excluded from shared-room activity. Detection is based on audio level with a one-second silence threshold; it is not a speech-recognition or identity-verification result.

Private conference coaching

Use <Conference coach="targetCallId">room</Conference> to let a supervisor hear the conversation while only the selected participant hears the supervisor. The target must be an active, unmuted, unheld member of the same tenant and room. To enable coaching or change targets during a call, update the supervisor participant with {"coaching":true,"callSidToCoach":"targetCallId"}; use {"coaching":false} to return to the shared conversation. Send coaching controls separately from mute/hold controls and wait for their completion receipts. A coach cannot be held. If the target leaves, is muted, or is held, private coaching stops before follow-up audio. Subscribe to conference modify callbacks for successful coaching changes.

Speech markup

Approved managed Polly accounts can nest break, say-as, sub, phoneme, p and s inside Say. For example: <Say voice="Polly.Joanna-Neural">Call <say-as interpret-as="telephone">2025550101</say-as>.<break time="250ms"/> Thank you.</Say> Existing text and audio duration limits still apply. Other tags and external audio references are rejected. The default lab voice and BYO synthesis v1 accept plain text only. Polly controls pronunciation; characters/spell-out on neural voices may use its related standard voice for that sentence.

Browser and native-cell audio

To check an existing active call without writing frontend code, sign in to your Kalla account, refresh Call activity with a key that permits Call history and Browser audio grants, then choose Get audio code. Open the audio page on your device and enter the private one-use code within ten minutes. This joins the selected call; it does not dial a new number. Revoking the key or ending the call invalidates unused codes. Use the browser SDK below for your application integration.

For browserSubject calls, your backend waits for browser.media_ready before issuing connect_destination. No customer or group member rings before the initiating browser is connected. Browser grants are passed to the Kalla audio client at https://media.kallavoice.com/rtc/, never an account key. The authenticated backend determines subject identity; a caller-supplied UID is not authentication.

For staff-cell calling, replace browserSubject with callerCell (cell.bridges permission). Staff must answer before Kalla connects the destination. Both numbers must be allowed by the tenant. Native-cell and browserSubject cannot be combined. Carrier limits remain operator-controlled during beta.

3. Invite someone or claim an incoming call

Create an invitation with POST /v1/calls/{call}/invitations: {"subject":"staff-123","purpose":"answer"}. Purpose is answer (one staff answer owner) or participant (add a person). Required permissions: calls.control, conferences and browser.grants. Invitations expire after 60 seconds unless joined.

GET /v1/invitations?subject=staff-123 lists the subject's invitations. GET /v1/invitations/{id} returns id, call_id, subject, state, version and expires. Both need calls.read.

POST /v1/invitations/{id}/decisions
{"subject":"staff-123","claimId":"device-attempt-unique-id","action":"accept","version":1}

Actions: accept, decline, busy, cancel. One device wins; another staff member's answer invite becomes answered_elsewhere. Cancellation removes the invited leg and releases its claim without ending the parent caller. A backend native-cell accept can additionally supply callerCell (cell.bridges required).

For an incoming browser screen, the backend POSTs an empty object to /v1/invitations/{id}/token. Give only that restricted token and invitation ID to https://media.kallavoice.com/rtc/incoming.html. It supports accept/decline/busy for that invitation, not arbitrary calling or account access. Do not embed the token in a URL.

Full saved-group orchestration and transfers remain under development. Native group joins require pressing 1; a carrier answer alone does not qualify. The private prompt waits up to eight seconds after playback. A wrong digit, timeout, revoked invitation, or ended parent prevents joining. This applies to add_participant and participant-purpose cell invitations; answering the original incoming call is separate.

4. Run a bounded call script

Dial defaults to 30 seconds ringing and four hours connected; carrier duration and spending limits may be stricter. Queue waits permit four hours. Script chains have a 24-hour budget inherited by replacements. Say/Play accept loop counts through 1,000; zero means up to 1,000 repetitions or until interrupted. Gather prompts currently require one repetition.

POST /v1/programs/validate with {"xml":"..."} checks the XML subset, not runtime eligibility. To run it, POST {"xml":"...","version":3} to /v1/calls/{call}/programs, with an Idempotency-Key. Requires calls.control and routing plus permissions for included operations. Poll GET /v1/programs/{id}. An empty <Redirect/> reloads the current document URL, including its query. Inline XML must supply a documentUrl to use it. The documented hop/deadline limits still apply. For application-driven redirection, POST {"version":3,"url":"https://your-app.example/next","method":"POST"} or {"version":3,"xml":"..."} to /v1/calls/{call}/redirect. It starts or replaces the script, retains an idempotent receipt and fetches a signed URL at most once per request. A failed/uncertain fetch is not replayed and leaves the existing script intact. To replace an active script, POST the same request shape to /v1/calls/{call}/programs/replace. The successor waits in replacement_wait until prior media cleanup is confirmed; uncertain cleanup blocks execution. Retries must use the same Idempotency-Key. The parent call stays connected.

<Response>
  <Say>Thank you for calling.</Say>
  <Hangup/>
</Response>

Implemented subset: Say, Play, Pause, Hangup, signed GET/POST Redirect, DTMF Gather, bounded Record, parallel/sequential Number dialing with private answer prompts, moderated conferences and Enqueue with XML waiting prompts, Leave and action continuation. A program ends the call when its instructions are exhausted. Unsupported behavior fails preflight.

Gather requires an HTTPS action callback or documentUrl on a registered tenant origin. DTMF and configured English speech input are supported. Dial supports Number, Client, mixed Number/Client, and approved SIP routes; parallel/sequential selection, private prompts, progress callbacks, and mixed recordings from answer or ringing. Conference and queue wait URLs accept bounded audio or XML. Dual recordings, incoming REFER, broader speech models, and some conference/SIP options remain incomplete. Outgoing Refer requires an explicitly approved SIP transfer route. See the specific sections below; full TwiML compatibility is not yet established.

GET /v1/recordings and GET /v1/recordings/{id}/audio require recordings.read. WAV download is available only after completion and remains tenant-scoped.

Live transcription

Configure a recognizer in Live transcription on your account, then validate it with synthetic silence. Credentials are encrypted and never returned. Changing or disabling the connection stops existing sessions; validation applies only to that configuration. The provider must implement kalla.asr.v1. An ordinary transcription HTTP API is not a compatible WebSocket service.

<Response>
  <Start><Transcription name="captions" track="both_tracks" partialResults="true" statusCallbackUrl="https://your-app.example/transcripts" /></Start>
  <Dial><Client>reception</Client></Dial>
  <Stop><Transcription name="captions" /></Stop>
</Response>

Supported attributes: name, statusCallbackUrl, languageCode, track, inboundTrackLabel, outboundTrackLabel and partialResults. Track defaults to both_tracks; partial results are off unless enabled. Provider/model choices belong to the connection. Unsupported provider-specific TwiML attributes are rejected. One transcription session per call; maximum four concurrent sessions on this worker. The connection sets a maximum duration and a daily track-audio budget, reserved before connecting. Two tracks reserve twice the session duration. Uncertain provider usage is not automatically refunded or retried.

Signed POST events are transcription-started, transcription-content, transcription-stopped and transcription-error. Content includes Track (inbound_track or outbound_track), Final, LanguageCode and JSON TranscriptionData with transcript and optional provider confidence. Labels are included in the start event. SequenceId increases within the session; consumers must deduplicate and handle callback retries. Inbound means audio received from the call's source channel; outbound means audio sent to it. Callback responses never control calls.

Start is nonblocking; transcription continues across script steps until Stop, hangup, expiry or failure. Audio buffering is bounded to 500ms per track. Provider failure stops transcription while retaining the call unless native audio capture cannot be confirmed stopped. Key/feature/connection revocation blocks and clears undelivered transcript notifications. A controller restart does not replay uncertain audio. Callback transcripts already delivered to your application remain your responsibility.

Connection API: GET/PUT/DELETE /v1/transcription/connection and POST /v1/transcription/connection/validate. PUT accepts endpoint (public wss://), credential, model, language, dailyAudioSeconds and maxSessionSeconds. Configuration requires intelligence.configure and intelligence.transcription; scripts additionally require calls.control, routing and media.streams.

The provider receives start with version kalla.asr.v1, requestId, track, model, language, encoding mulaw and sampleRate 8000. It returns ready with the same version/requestId, then receives 160-byte binary audio frames. Results have event result, consecutive sequence beginning 1, boolean final, text and optional confidence. Kalla acknowledges persisted results with ack/sequence. On finish, the provider drains results and returns finished with the last sequence. Every JSON envelope carries version and requestId. Control messages are limited to 16KiB; text to 4096 characters. Final drain is limited to five seconds. No automatic reconnect or audio replay occurs.

Answer without a warm application server

In Application settings, leave Incoming-call URL empty and save Hosted incoming-call TwiML. Kalla validates and stores the document, then executes it locally on each incoming call. You can also PUT /v1/routing/inbound-script with version, url (empty), method (POST), and xml. GET returns the current version and document; stale updates fail. Choose either a webhook URL or hosted XML, never both. Clear both to restore default routing.

Call-event delivery is asynchronous and does not hold up hosted call instructions. Hosted scripts may still depend on external services: Redirect, action URLs, remote audio, speech providers and dynamic lookups can add latency. Referenced callback origins must already be registered before saving the script. Keys, features, approved destinations and quotas are checked again at call time. This does not cache authentication or customer data.

Hosting the initial instructions removes the initial TwiML webhook cold start. Replacing an application's answer-claim and conference endpoints requires its integration to use Kalla's own controls; changing this setting alone does not remove every backend dependency.

Dial a queue

<Dial><Queue>sales</Queue></Dial> connects to the oldest eligible scripted Enqueue caller in the tenant’s configured queue. An empty queue waits up to Dial timeout. Each caller can be claimed by only one dequeuer. An optional Queue url/method runs Say, Play, Pause or Redirect privately on the queued caller before connection; it receives QueueSid, QueueTime and DequeingCallSid. Mixed recording belongs to the queued caller’s CallSid. Cancelling or replacing either script drains shared audio before continuing. Controller restart ends both bound calls instead of guessing bridge ownership. Queue TaskRouter reservation options and speech waiting menus remain unsupported.

Conference control API

GET /v1/conferences supports optional name and state filters (waiting, in-progress, completed). Use the returned id with GET /v1/conferences/{id} or GET /v1/conferences/{id}/participants. Reads require calls.read and conferences. DELETE /v1/conferences/{id}/participants/{callId} hangs up only that participant. POST /v1/conferences/{id}/end schedules hangup for all remaining participants atomically. Writes additionally require calls.control and an Idempotency-Key; poll the returned command receipts. Completed conference history remains queryable. To mute or hold a joined participant, POST {"muted":true,"hold":true} (either or both booleans) to /v1/conferences/{id}/participants/{callId}. Poll /v1/conference-controls/{id} for completed, rejected or uncertain. Hold moves the participant into a private bridge with the server music-on-hold class; false resumes the same conference. A partial native failure disconnects only that participant. Custom hold URLs and coaching remain unsupported. To invite a new phone participant, POST to /v1/conferences/{id}/participants with {"to":"+12025550101","from":"+12025550102","label":"staff-123"}. Requires calls.create and routing as well. By default the called person privately presses 1 before joining; confirm:false explicitly disables that gate. Optional beep, muted, startConferenceOnEnter and endConferenceOnExit use conference semantics. The invitation is bound to this conference id; a late answer cannot join a replacement room. Pending invitations reserve labels and capacity. Only tenant-authorized numbers/carriers can be dialed.

Outbound answer scripts

Add "program":{"xml":"<Response><Say>Hello</Say></Response>"} to POST /v1/calls, or use "program":{"url":"https://your-app.example/answer","method":"POST"}. Requires calls.control and routing in addition to calls.create. The signed URL must use a registered callback origin. Inline XML is fully checked before accepting the call; execution starts after answer. Invalid or uncertain answer hooks end the call and are not retried. This option is for direct outbound calls, not browserSubject/callerCell calls. Direct calls accept an optional timeout of 1–60 seconds for ringing, and also accept statusCallback, statusCallbackMethod and statusCallbackEvent at the top level. These signed notifications use the API call id as CallSid and Direction outbound-api. Conference invitations accept the same callback options.

Child call progress

Number supports statusCallback, statusCallbackMethod (GET/POST), and statusCallbackEvent (space-separated initiated, ringing, answered, completed; default completed). Signed notifications report observed native state, stable CallSid, ParentCallSid, SequenceNumber and terminal CallDuration. Delivery can retry; deduplicate by X-Phone-Event-Id. GET /v1/calls/{call}/legs lists children and GET /v1/legs/{id} reads one. DELETE /v1/legs/{id} cancels only that leg; parent scripting continues. Requires calls.read and, for cancellation, calls.control. Unconfirmed native cleanup returns an error, not success.

Inbound routing scripts

GET or PUT /v1/routing/inbound-script. PUT accepts {"version":0,"url":"https://your-app.example/incoming","method":"POST"}. The HTTPS origin must already be registered for tenant callbacks. Configuration requires routing, calls.control and voicemail. Set url to an empty string to disable the hook. Kalla evaluates the signed hook before answering; a sole Reject verb rejects the call without answering. Other supported scripts execute after answer. Hook failures use cloud voicemail. Revoking the configuring key disables its authority; configure with a long-lived server key.

Bring your speech service

Integration preview: a tenant can choose managed Polly or connect its own synthesis endpoint for Say. This is distinct from an AI conversation connection: Portable Mind owns its recognition, voice, turn-taking and mind; Kalla supplies telephone transport. Walkie connects directly to Portable Mind.

When enabled by the operator, use PUT /v1/speech/connection with a backend API key carrying speech.configure. GET reads redacted status; POST to /v1/speech/connection/probe validates a synthetic prompt before activation; DELETE disables it. Credentials are encrypted and never returned. Keep backend API keys out of browsers.

{"provider":"byo","voice":"my-voice","language":"en-US",
 "endpoint":"https://speech.example.com/v1/synthesize",
 "credential":"YOUR_ENDPOINT_SECRET","dailyCharacters":10000}

Kalla POSTs schemaVersion 1, requestId, text, voice, language, audio {encoding: pcm_s16le, sampleRate: 16000, channels: 1}, and maxAudioSeconds: 30. Authentication is Bearer with an Idempotency-Key. Respond with HTTP 200, application/octet-stream, X-Kalla-Request-Id matching the request, and raw 16 kHz mono PCM (not WAV), up to 30 seconds. The deadline is eight seconds. Redirects, compressed responses and private-network destinations are rejected unless the operator explicitly enrolls that exact private endpoint.

After validation, Say without an explicit voice uses the selection. Explicit voices/languages must match. Disabled or failed connections never fall back to another provider. Quotas include uncertain attempts and probes. Managed Polly uses provider polly, the configured catalog voice, and empty endpoint/credential fields. Managed speech pricing and invoicing are not active yet; do not interpret a successful probe as commercial availability. Existing cloud calling does not depend on a customer's AI engine remaining online.

Live audio stream controls

Use POST /v1/calls/{call}/streams with an Idempotency-Key and {"url":"wss://voice.example.com/audio","name":"observer","track":"both_tracks","parameters":{"purpose":"analysis"}}. The endpoint must belong to a registered callback/audio origin. Required permissions: calls.read, calls.control, routing and media.streams. Optional statusCallback and statusCallbackMethod use signed callbacks.

GET the collection or /streams/{id-or-name} for state. Creation returns 202 with a stable stream ID; starting is not proof of a connected WebSocket. Poll for active. POST {"status":"stopped"} to the individual stream to request cleanup, then poll for stopped. Reusing the create Idempotency-Key cannot reopen a stopped stream. The parent call and its script keep running.

REST creation is unidirectional. A bidirectional Connect stream must be interrupted through the existing call/program control, not this stop endpoint. These are Kalla v1 endpoints, not Twilio's raw account-SID REST URLs. The server SDK supports client.calls(callId).streams.create(...), .streams.list(), and .streams(name).update({status: 'stopped'}). Stream states also expose starting, failed and uncertain rather than hiding pending outcomes.

Configure your provider, selected outputs, daily job limit and automatic processing from Recording analysis in the account console. Automatic processing is opt-in and applies to new recordings only. Provider credentials are encrypted and never included in downloaded application configuration. Results remain available through the API below.

5. Optional recording intelligence

Separate features: transcription, keywords, tone, score and summary. Enable each feature and grant its corresponding intelligence.* permission. These are assessments, not authoritative facts about a person. No model output executes a CRM action.

POST /v1/intelligence/profiles requires intelligence.configure and each requested feature. Exact fields: provider (gemini or self-hosted), model, credential, endpoint, features, language, rubric, maxJobsPerDay (1–100), cloudEgressAllowed and automatic. A score requires a rubric. Provider credentials are encrypted and never returned.

{
  "provider":"self-hosted", "model":"your-model",
  "credential":"YOUR_PROVIDER_SECRET", "endpoint":"https://your-service.example/analyze",
  "features":["transcription","summary"], "language":"en",
  "rubric":"", "maxJobsPerDay":10, "cloudEgressAllowed":false, "automatic":false
}

Automatic mode queues new completed recordings after activation. No silent cloud fallback. Self-hosted inference must implement Kalla's JSON audio contract; private endpoints require operator registration. Gemini requires explicit cloud egress consent. No model quality or live pricing is guaranteed in beta.

POST /v1/intelligence/jobs with recordingId and profileId; GET or DELETE /v1/intelligence/jobs/{id}. Delete clears derived output while retaining a tombstone. DELETE /v1/intelligence/profiles/{id} disables future processing. Reads require intelligence.read, recordings.read and the profile's feature permissions. Deletion/configuration additionally requires intelligence.configure. In-flight provider requests cannot be recalled. Jobs with uncertain outcomes are never automatically resubmitted.

6. Events and webhooks

GET /v1/events?after=0 requires events.read. Returns events and nextCursor. Persist your cursor. Events include id, sequence, callId, type, data, occurredAt and schemaVersion. Use event IDs to deduplicate; late events must not reopen a finished call.

Webhook registration is operator-managed during beta. Signed JSON delivery is at least once. Headers: X-Phone-Event-Id, X-Phone-Timestamp, X-Phone-Signature. Verify HMAC-SHA256 over timestamp + "." + raw_request_body using your webhook secret, compare in constant time, reject timestamps older/newer than five minutes, and deduplicate event IDs. Acknowledge only after durable receipt. Do not parse and reserialize JSON before verification.

Script actions support GET query parameters or POST form data with X-Phone-Request-Id and the same timestamp/signature construction. Actions must return bounded XML. GET signatures cover the URL-encoded full query (original query pairs followed by call fields); POST signatures cover the body. Status notifications accept any 2xx, are retried up to ten attempts, and include a stable X-Phone-Event-Id for deduplication. Notifications may arrive out of order; use conference SequenceNumber. Redirects and private-network destinations are rejected.

Compatibility additions

Empty and exhausted Responses end the call without an explicit Hangup. Program creation accepts an optional documentUrl on a registered callback origin; relative links and omitted Gather/Record actions resolve against it. Inline Gather still needs an explicit action when no documentUrl is supplied. Play digits supports keypad tones and w/W half-second/one-second waits, but cannot be nested in Gather. Number sendDigits also supports waits. Play accepts bounded WAV, MP3, AIFF, GSM and μ-law audio through a restricted decoder; the 30-second prompt limit remains.

Conference statusCallback supports start/end/join/leave and GET/POST; the first participant chooses the callback and events. A conference starts after at least two call sessions join and a starter is present. Record recordingStatusCallback supports in-progress/completed/absent with an authenticated recording URL. Number url accepts private Say/Play/Pause/Gather/Redirect/Hangup instructions before bridging, with a 60-second deadline; exhaustion admits the leg and explicit Hangup declines it. Record supports trim-silence (default), do-not-trim, finish-key sets and timeout=0. Edge trimming retains 100 ms around detected audio and preserves wholly silent recordings. Recording callbacks use Kalla IDs and signatures; existing Twilio webhook handlers need the provider adapter. Full TwiML and GSE CRM parity are still in progress.

Dialing and conferences

Dial accepts up to ten distinct allowlisted Number targets. It also accepts up to ten Client identities in a separate client-only Dial: the first acceptance cancels other invitations. Client dialing supports nested Identity and Parameter elements for caller context. Custom parameters are delivered to the authorized invitation receiver and do not grant permissions. Client url supports private Say/Play/Pause/Redirect/Hangup instructions before shared audio. Client statusCallback supports initiated, ringing, answered and completed with a stable child call ID. Cancel that child through DELETE /v1/legs/{id}; the parent stays active. Private Gather is supported for Number and Client answer URLs, including mixed dialing; sequential Client and mixed Number/Client dialing are supported. By default it rings simultaneously, connects the first answered leg and cancels the others. Set sequential="true" to try targets in document order, stopping after the first answer. The full sequence of ringing windows counts toward the 24-hour script deadline. Paid-leg capacity is reserved for all targets before the first ring; failed or cancelled attempts are not refunded automatically. timeout defaults to 30 seconds and timeLimit defaults to 14,400 seconds. answerOnBridge accepts true or false (default). A first Dial with true defers the incoming answer until its destination, browser, queue pickup, or conference is admitted. Number sendDigits accepts up to 32 keypad digits (0–9, * and #). This is first-answer dialing, not staff group admission.

<Response>
  <Dial timeout="20" timeLimit="120" answerOnBridge="true" sequential="true">
    <Number>+12025550101</Number>
    <Number sendDigits="123#">+12025550103</Number>
  </Dial>
  <Hangup/>
</Response>

Numbers above are examples; your tenant and activated carrier must permit every target and caller ID.

A Conference must be the only Dial child. Room names are scoped to your tenant. Supported attributes: startConferenceOnEnter, endConferenceOnExit, muted, maxParticipants (2–25), participantLabel (unique per room), and beep (false, true, onEnter, onExit). Dial and Conference recordings default to trim-silence; use trim="do-not-trim" to retain exact timing. Conference Dial actions omit DialCallSid and DialCallDuration. Beeps default to true; enabled tones last at most 200 milliseconds and are heard by the admitted room. An exit tone is omitted when the room ends or is empty. Members wait in their existing private call until a starter arrives. A member with endConferenceOnExit ends the room on leaving; the other scripts continue. Capacity counts joined call sessions, which may already contain multiple audio legs. Attach browser audio before starting the program. Current worker/call capacity may be lower than maxParticipants. A controller restart ends conference calls rather than reconstructing uncertain membership.

<Response>
  <Dial timeLimit="120">
    <Conference startConferenceOnEnter="true" endConferenceOnExit="true"
      maxParticipants="8" participantLabel="moderator" beep="false">Sales</Conference>
  </Dial>
  <Hangup/>
</Response>

Bring your own carrier

In your account, enroll a TLS/SRTP SIP connection using IP authentication, digest credentials without registration, or outbound registration. Secrets are encrypted and are never returned in account status. Supply your provider's public IPv4 source networks and a SIP hostname with a valid TLS certificate.

  1. Enter inbound numbers, explicit outbound destinations, and requested call-count/duration limits. Site requests expire after seven days and do not authorize spending.
  2. Run Check DNS and TLS. Private DNS answers are rejected. A successful certificate check confirms transport reachability only.
  3. An operator verifies SIP authentication, SRTP media and number ownership, reviews your requested budget, and installs the private configuration bundle with pinned DNS, source ACLs and per-connection routing authentication.
  4. Activation requires the matching installation receipt and fresh validation. Until then, no carrier route is enabled. Suspend connection stops new inbound and outbound calls; it does not interrupt an existing call.

The provider must support the private X-Kalla-Route header on incoming calls. Shared provider IP ranges alone never select a tenant. Duplicate DID claims are rejected. Activation and budget approval are operator actions, unavailable to customer API keys. No number is ported or rerouted automatically. Existing test usage is preserved when a legacy test trunk is adopted. Readiness shows policy expiry, payment requirements and remaining approved attempts; exhausting outbound attempts does not disable otherwise valid incoming calls. Repeated private exports preserve their manifest and route secret until configuration or validated network details change. Providers without TLS/SRTP or the required inbound header are not supported by this release.

Server-side Node adapter

Download the Kalla client and callback verifier. The client covers calls and conferences with a familiar provider interface. Use your API key only in a trusted server process. Unsupported options fail explicitly; accepted changes include a receipt to check for actual completion. It does not replace browser media libraries or generate Twilio tokens.

const {KallaClient} = require('./kalla-client.cjs');
const kalla = new KallaClient({apiKey: process.env.KALLA_API_KEY});
const call = await kalla.calls.create({
  to: '+12025550101', from: '+12025550102',
  twiml: '<Response><Say>Hello</Say></Response>'
}, {idempotencyKey: 'your-stable-operation-id'});

Lost replies never cause an automatic redial. Preserve the operation id and inspect its outcome before deciding to retry. Server script URLs must use a registered callback origin.

Callback signatures

New integrations must verify X-Phone-Signature-V2. Its HMAC-SHA256 input is compact UTF-8 JSON [timestamp,method,targetUrl,identity,kind], a period, then the authenticated payload bytes. kind is event or request; identity comes from the corresponding X-Phone-Event-Id or X-Phone-Request-Id. The target URL is X-Phone-Callback-Target and must match the receiver’s configured public origin/path and original query parameters. POST signs raw body bytes; GET signs canonical form-encoded query pairs, including the original query. Accept timestamps within five minutes and persist event IDs with side effects to deduplicate retries. The server-side adapter includes a verifier. V1 headers remain for older test consumers but do not bind the destination; use V2 for new integrations.

Ending a conference

POST /v1/conferences/{id}/end returns an accepted receipt. Poll GET /v1/conference-ends/{receiptId} for completed, rejected or uncertain. Ending a conference lets each call continue after Dial or execute its action callback; it does not substitute call hangups. Use the call hangup command to explicitly end a call. Pending conference ends block new participant invitations.

Waiting callers

Queue waitUrl supports keypad Gather alongside Say, Play, Pause, Redirect, Leave and Hangup. A Gather action receives Digits and current queue fields; leaving resumes the Enqueue action without ending the call. Queue and conference wait URLs may also return bounded WAV or MP3 directly. Set waitUrl to an empty string for silence. Prompts and digit collection stop before pickup or room admission. Speech and mixed speech/keypad input are available when the operator has configured the private recognizer and the key includes speech.input. The beta supports en-US with vosk-small-en-us-0.15; accuracy is still under evaluation.

Recording downloads

GET /v1/recordings/{id} returns tenant-scoped metadata. GET /v1/recordings/{id}/audio and /audio.wav return WAV; /audio.mp3 returns MP3. All require recordings.read and a completed recording. Use your authenticated backend to proxy playback; API keys must never appear in public audio URLs. The Node adapter exposes recordings.list, fetch and download with an explicit wav or mp3 format.

Long recordings

Record defaults to one hour and supports up to four hours. Dial/Conference recordings support four hours. Authenticated WAV and MP3 endpoints accept single byte ranges for seeking. Downloads use private disk-backed exports, with a 1 GiB input cap and two simultaneous exports. Use the Node SDK downloadStream({format: "mp3"}) for incremental retrieval; download() returns a Buffer capped at 64 MiB. Recording analysis has separate limits.

Browser keypad

The browser audio client exposes sendDigit("1") for an active Kalla Gather prompt. A connection-scoped token is returned with the media answer; it cannot control another call. Duplicate request IDs do not repeat a digit. A private Client answer prompt can collect keypad input before admitting that participant. External carrier IVR/RTP DTMF is not supported by this control path yet.

Media streams and speech input

Select media.streams or speech.input in your feature and API-key settings, alongside calls.control and routing. These features are isolated beta capabilities.

<Start><Stream name="listen" url="wss://your-app.example/audio" track="both_tracks"/></Start> forks caller audio while the script continues. Stop a named unidirectional stream with <Stop><Stream name="listen"/></Stop>. Connect/Stream blocks until the socket closes and supports inbound audio plus paced PCMU/8000 playback, marks and clear. Register the matching HTTPS callback origin first. Only public WSS/443 is allowed; query strings and redirects are rejected. The handshake uses Kalla V2 signatures, not Twilio authentication. Limits are four tracks per call, eight stream sessions in this lab, five seconds of queued playback and a four-hour maximum. Raw audio is not logged or saved.

Gather accepts input="speech" or input="dtmf speech", language="en-US", speechModel="vosk-small-en-us-0.15", and speechTimeout="auto" or 1–30 seconds. The operator must configure the local recognition service. Speech results use SpeechResult and optional Confidence; keypad results use Digits. Signed partialResultCallback notifications carry UnstableSpeechResult. Each collection is bounded to 60 seconds. Other models, languages and speech hints are not supported yet. The small model has not established real-caller accuracy.

Recording transcription

Use <Record transcribe="true" transcribeCallback="https://your-app.example/transcript"/> with your configured analysis provider. The key needs recordings.read, intelligence.read, intelligence.transcription and the configured provider's other feature scopes. The existing provider quota applies; Kalla never chooses a substitute provider. maxLength defaults to 120 seconds for transcription. Recordings longer than two and shorter than 120 seconds are eligible. Missing or ineligible audio is skipped without provider dispatch or callback.

Results are asynchronous and never change call flow. Signed POST callbacks include TranscriptionSid, TranscriptionStatus, TranscriptionText, TranscriptionUrl, RecordingSid and CallSid. Without a callback, fetch recording metadata for transcription.jobId and state, then retrieve the analysis job through the authenticated API. Provider uncertainty is reported as failed and is never automatically retried. Revoking the requesting key or provider blocks pending transcript disclosure. Audio stays on the phone server unless your configured analysis provider requires an upload. Provider configuration and explicit cloud-egress consent govern that upload.

Conference control events

Conference statusCallbackEvent accepts start, end, join, leave, mute and hold. Selecting mute includes participant-mute and participant-unmute; hold includes participant-hold and participant-unhold. Notifications share the room generation's sequence, include the participant's current Muted/Hold values, and are sent only after successful native control. Repeating an unchanged setting does not create a transition event. The first participant chooses the room's callback subscription.

SIP transfers

<Refer action="https://your-app.example/transfer"><Sip>sip:approved@example.com</Sip></Refer> transfers an answered SIP leg only when the operator has explicitly enabled transfers for that tenant's exact destination. Ordinary SIP route approval does not enable transfers. Transfer requests consume the route's durable attempt limit; uncertain outcomes are never automatically repeated.

The signed action includes ReferSipResponseCode for the observed response to REFER, and NotifySipResponseCode only when the peer supplied a SIP NOTIFY result. ReferCallStatus reports in-progress after a successful NOTIFY, busy for a busy destination, or failed for a refusal or destination error. An accepted REFER with no NOTIFY does not prove that the new call connected, so that call-status field is omitted. KallaTransferStatus (SUCCESS, FAILURE or UNSUPPORTED) and KallaTransferResponseCode remain available for compatibility. Codes are captured separately for each transfer; an earlier result is never reused. A missing result ends the owned call after a bounded wait. Refer URI X-headers are validated and percent-encoded in Refer-To; they do not change the tenant-approved base destination or bypass transfer budgets. Routing/control headers and CR/LF are rejected. Incoming REFER and transfers on non-SIP legs remain unsupported. A pending transfer cannot be replaced by another script; hangup remains available.

Errors and operational limits

Error bodies are {"error":"machine_readable_code"}.

StatusMeaning
400Invalid or unsupported request. Correct it before retrying.
401 / 403Invalid key, insufficient feature/scope, wrong origin or disallowed destination.
404Resource absent from the authenticated tenant.
409State/version conflict, request already claimed, unavailable feature or capacity/budget limit. Inspect before retrying.
429Rate or quota limit. Back off; do not create replacement operation IDs.

The default platform capacity is 64 active calls across tenants; individual tenant limits still apply. At capacity, new inbound/outbound admissions fail before any dial. Long scripts use bounded concurrent dispatch with reserved hangup/control capacity. General API requests are limited to 16 KiB. Recording analysis accepts bounded completed WAV files, not arbitrary remote audio URLs. Beta enrollment does not activate public carrier routing or live billing. Maintain caller consent and your approved recording policy. Emergency calling is not enabled by this beta.

Conference hold audio

Participant updates accept hold, muted, and an optional holdUrl with holdMethod (GET or POST). For example: {"hold":true,"holdUrl":"https://your-app.example/hold","holdMethod":"POST"}. Register that HTTPS origin in Application settings first. The signed request includes the call and conference identifiers. Return bounded Say/Play/Pause/Redirect XML or WAV/MP3 audio. Completed documents repeat while held. Audio plays only in the held participant’s private bridge and is stopped before rejoining the conference. Send {"hold":false} to resume. Fetch or playback failure ends the affected leg rather than exposing conference audio. Without holdUrl, the built-in hold music is used.

Managed speech fees

When managed Polly pricing is published, the account form shows the price version, eligible voices and fee per provider-confirmed character. Accept that fee before saving. API clients include the catalog’s pricing.fingerprint as priceFingerprint in the speech connection. A new price version requires new acceptance. Validation phrases also consume provider characters. Speech has a separate usage meter from calling; bring-your-own provider fees are paid directly to that provider. Live speech billing is not yet activated.

Recording history

Find recordings for a specific call with GET /v1/recordings?callId=CALL_ID&limit=100. Filtering happens before pagination, so older calls remain searchable. The response includes recordings and nextPageToken; pass that token as pageToken with the same call filter to fetch the next page. Page size is 1–100. Each request requires your recording-read permission.

const page = await client.recordings.page({
  callSid, pageSize: 100, pageToken
});
// page.records, page.nextPageToken

Recording tokens do not authorize access. API keys remain server-side.

Choose a scripted recording track

Dial supports recordingTrack="inbound" for audio received from the caller, recordingTrack="outbound" for audio sent to the caller, or recordingTrack="both" for the default mixed recording. Directional files are mono and their metadata and signed callbacks identify the selected track. The same recording-read permissions protect every file. A recording requested on a parent Dial also follows that caller through a Conference, including directional or dual audio. It ends when that caller leaves; a separate record-from-start on the Conference continues for the room. Shared-room recordings are mixed mono. For separate caller/outgoing channels, choose record="record-from-answer-dual" or record="record-from-ringing-dual". The left channel contains caller audio; the right contains audio sent to the caller. Selecting only inbound or outbound still produces a mono file. API recording creation also accepts recordingChannels: "dual". Both channels obey recording permissions and silence/skip pause controls; a failed half prevents publication of an incomplete pair.

<Dial record="record-from-answer" recordingTrack="inbound">
  <Number>+12025550101</Number>
</Dial>

Legacy Dial record="true" and record="false" select recording from answer and no recording respectively.

Answer timing

For an incoming call whose first instruction is Dial, answerOnBridge="true" keeps the caller ringing until the winning destination has answered and completed its private answer URL. Call status remains ringing until the native answer succeeds. The default and answerOnBridge="false" answer before dialing. For Queue, answer waits for pickup and its private URL; for Conference, answer waits until the room admits the participant. Existing answered calls stay answered.

Regional ringback

Dial Number, Sip, Client, or mixed Number/Client targets can set ringTone="us". Ringback plays only to the caller during dialing and private destination preparation. It stops and drains before admitting the winner. Supported regional tones: at, au, bg, br, be, ch, cl, cn, cz, de, dk, ee, es, fi, fr, gr, hu, il, in, it, lt, jp, mx, my, nl, no, nz, ph, pl, pt, ru, se, sg, th, uk, us, us-old, tw, ve, za. Omit this option to preserve the existing call's signaling/audio behavior. Regional tones do not determine routing or carrier location.

Reusable recording policies

Create immutable policies in the customer console or POST /v1/recording-configurations. GET lists your policies; DELETE /v1/recording-configurations/{id} disables future use. Requires Routing and Voicemail permissions. Existing accepted scripts retain their resolved options; disabling does not stop an existing recording.

{"name":"Conference archive","compositionPolicy":{"channels":"mono","track":"both","trim":"do-not-trim"},"conferenceRecordingStatusCallback":{"url":"https://your-app.example.com/recordings","method":"POST","events":["completed"]}}

Use the returned ID as recordingConfigurationId on Dial, Record, Conference, or in the call recording start API. Do not combine it with explicit trim, track, channels or recording callback options. Callback origins must already be registered. Policies support separate callRecordingStatusCallback and conferenceRecordingStatusCallback objects, GET/POST and in-progress/completed/absent events.

<Response><Dial><Conference recordingConfigurationId="RC...">support</Conference></Dial></Response>

Composition defaults: dual channels, both tracks, do-not-trim. For Record, select mono/both. Conference supports mono/both or dual/both: dual isolates the first admitted participant from the remaining participants and continues after that first participant leaves. Set recordingChannels="dual" directly on Conference or choose a dual policy. Mono remains the default. Muted, held and coaching participants are excluded from dual capture; room-wide announcements and beeps are included on the second track. Private announcements stay private. Silence/skip pauses, resume and stop apply to both tracks. Dial and call recording API support directional dual capture. Provider-specific postprocessing features are rejected; use Kalla’s separate recording analysis integration.

Secure SIPREC recording connector

Configure your recorder in the customer console or use GET/PUT/DELETE /v1/siprec/connection and POST /v1/siprec/connection/validate. Configuration requires routing, media.streams and voicemail permissions. The configuring key must stay active. Credentials are encrypted and never returned. Each change requires a successful synthetic validation before use. Disabling the connection stops its recordings while preserving healthy calls.

{"name":"company-recorder","endpoint":"sips:recorder@srs.example.com:5061","credentialHeader":"X-Auth-Token","credential":"your-secret","mediaNetworks":["your-public-media-IP/32"],"dailyAudioSeconds":3600,"maxSessionSeconds":600}

Use real public IPv4 media networks, at most16 ranges, each /24 or narrower. This profile requires TLS1.2 or later with a publicly trusted certificate matching the configured hostname, a custom X- authentication header, PCMU/8000 audio and SDES-SRTP AES_CM_128_HMAC_SHA1_80. Your SRS must use the established TLS connection for dialog requests. Redirects, cross-host Contacts, Record-Route proxies, cleartext media, digest authentication and IPv6 media are not supported. No insecure fallback occurs. The SRS owns recording storage and processing.

<Response>
  <Start><Siprec name="archive" connectorName="company-recorder" track="both_tracks" statusCallback="https://your-app.example.com/recording-events"/></Start>
  <Pause length="30"/>
  <Stop><Siprec name="archive"/></Stop>
</Response>

Start is asynchronous and requires calls.control, routing, media.streams and voicemail. Continue the call with another instruction. Default track is inbound_track; outbound_track and both_tracks are also supported. One active SIPREC session per call, four per controller. Custom Parameter elements are rejected in this profile. Metadata uses opaque participant identifiers, not customer names or numbers. Audio is a read-only fork; recorder audio cannot enter the call.

Callbacks use your registered HTTPS origin and Kalla’s standard signatures. Fields include AccountSid, CallSid, SiprecSid, SiprecName, SiprecEvent, SequenceId and Timestamp. Events are siprec-started, siprec-stopped and siprec-error; errors add SiprecError and SiprecErrorCode. GET or POST is supported. Deduplicate by notification ID. Pending notifications are blocked and scrubbed when their originating authority is revoked.

Quota reserves the maximum session duration for each selected track before connection. Confirmed completion settles to rounded-up audio seconds; uncertain outcomes retain the reservation. Validation reserves two seconds and sends half a second of synthetic silence per track. Sessions never silently reconnect. Call hangup, explicit Stop, permission revocation and controller restart stop native capture. An unconfirmed capture cleanup ends only its owned call. Caller recording notices and consent remain your application’s responsibility.