Detect scam calls using Go with the LLM Gateway and Twilio
Build a Go service that transcribes Twilio calls in real time and flags likely scam attempts with the AssemblyAI LLM Gateway. Full working code.



Twilio will happily give you a live copy of every call your number handles. That's what Media Streams does: it forks the call audio to a WebSocket you control, in real time, while the call is still going. What it won't give you is words — you get raw mulaw frames at 8kHz and it's your problem from there.
So the interesting question isn't "can I get the audio." It's what you do with it inside the two seconds before the caller finishes their sentence.
This tutorial builds the whole pipeline in Go: a Twilio number that forks its audio to your service, real-time transcription over a WebSocket, a running transcript held in memory, and an LLM pass at the end of the call that reads the transcript and decides whether it looks like a scam. Roughly 250 lines of Go, no framework, no SDK required. Everything here is complete and runnable.
Scam classification is the example because it exercises the hard parts — you need accurate entity capture (account numbers, dollar amounts, callback numbers), you need it live, and you need a model reasoning over the result. Swap the prompt and the same service does compliance monitoring, QA scoring, or intent routing.
What you'll build
Two WebSockets, one HTTP server. Twilio talks to you; you talk to AssemblyAI. The Go process sits in the middle holding state for the duration of the call.
Prerequisites
- Go 1.21 or newer.
- A Twilio account and a voice-capable phone number. Trial accounts work.
- An AssemblyAI API key. Sign up free — the key covers both streaming transcription and the LLM Gateway.
- ngrok, or any tunnel that gives you a public HTTPS/WSS URL. Twilio has to reach your laptop.
One dependency:
Why real-time and not after the fact
You could record the call, transcribe it when it ends, and classify the recording. It's simpler and it's the wrong shape for this problem.
A scam call is a live event. The value of knowing "this caller is walking your user through a gift card purchase" collapses to nearly zero once the call is over and the gift cards are bought. Everything useful — warning the user mid-call, flagging the session for a supervisor, cutting the call off, dropping a compliance marker into the CRM — depends on having the text while the audio is still arriving.
The architecture also composes. Once transcript text is flowing through your process turn by turn, you can act on any turn, not just the last one. This tutorial classifies at the end of the call because that's the simplest version to reason about, but the buffer it builds is exactly what you'd tap for mid-call intervention, and the same pipeline is the front half of any voice agent you'd build on the same number.
Step 1: Get a public URL with ngrok
Twilio needs to reach your machine from the public internet, both for the TwiML webhook and for the media WebSocket. Start the tunnel first so you have the hostname before you configure anything.
You'll get a forwarding line like:
Two URLs come out of that hostname:
- Voice webhook: https://a1b2-73-15-201-4.ngrok-free.app/voice
- Media stream: wss://a1b2-73-15-201-4.ngrok-free.app/stream
Note the scheme change — https for the webhook, wss for the stream, same host. Keep this terminal open; a free ngrok hostname changes every restart and you'll have to update Twilio when it does.
Export it for the server to build TwiML with:
That last variable is deliberately not hardcoded. The LLM Gateway fronts models from OpenAI, Anthropic, Google, and others behind one API and one bill, and the right model for a classification task like this changes as new ones ship. Pick one from the gateway's model list and set it here rather than pinning a model id into your source.
Step 2: Point a Twilio number at your server
In the Twilio Console, open Phone Numbers → Manage → Active numbers and click your number. Under Voice Configuration, set:
- A call comes in: Webhook
- URL: https://a1b2-73-15-201-4.ngrok-free.app/voice
- HTTP method: POST
Save it. From now on, every inbound call to that number makes Twilio POST to /voice and execute whatever TwiML you return.
The TwiML we'll return uses <Start><Stream>, which forks a copy of the inbound audio to your WebSocket and then keeps going with the rest of the document. That's the distinction that matters: <Connect><Stream> hands the call over to your socket entirely, while <Start><Stream> taps it and lets the call proceed. We want the tap.
The <Pause> keeps the call alive so there's something to listen to. In a real deployment this is where you'd <Dial> the agent, the queue, or the voice agent that actually handles the call.
What Twilio sends you
Once the stream opens, Twilio sends JSON text frames over the WebSocket. Four event types matter:
The audio format is the part people get wrong. Twilio Media Streams are PCMU (G.711 mu-law) at 8000 Hz, mono, base64-encoded, in 20ms frames — 160 bytes of mulaw per frame. Don't resample it, don't convert it to PCM16, don't assume 16kHz. Decode the base64 and forward the bytes exactly as they arrived, then tell AssemblyAI that's what's coming.
One thing you cannot skip: batch the frames
Twilio emits a media frame every 20 milliseconds. The AssemblyAI streaming API requires between 50 ms and 1000 ms of audio per WebSocket message. Forward Twilio's frames one-for-one and the very first one is rejected:
The socket then closes and you never receive a single transcript. This is the most common way this integration fails, and it fails quietly in the sense that your Twilio side looks perfectly healthy — audio is flowing, frames are arriving, and the only sign of trouble is one error frame before the connection drops.
The fix is a small buffer. Accumulate five Twilio frames (5 × 20 ms = 100 ms, 800 bytes of mu-law) and send that. 100 ms sits comfortably inside the window, adds negligible latency next to the model's own turn detection, and cuts your WebSocket write rate by 5×. The transcriber below does this internally so the rest of your code doesn't have to think about it.
Step 3: The real-time transcriber
Here's the transcriber. It opens a WebSocket to wss://streaming.assemblyai.com/v3/ws, declares the Twilio audio format in the query string, batches incoming frames to a legal message size, and dispatches turn events to callbacks.
Create transcriber.go:
A few notes on the choices in there.
encoding=pcm_mulaw and sample_rate=8000 are doing real work. The streaming API accepts mulaw natively, which means zero transcoding in your hot path — the bytes Twilio hands you are the bytes you write to the socket. Get either parameter wrong and you won't get an error; you'll get a transcript that reads like static.
The batching in SendAudio is load-bearing, not an optimization. Without it the session dies on the first frame with error 3007. Keeping the buffer inside the transcriber means the Twilio handler stays a straight decode-and-forward loop, and it means Flush on close doesn't lose the tail of the call.
speech_model=universal-3-5-pro selects Universal-3.5 Pro Realtime, the current streaming flagship. Telephony audio at 8kHz is the hardest input a speech model gets — it's band-limited, often compressed twice, and frequently noisy — which is exactly the case where model choice shows up in the output. On the Pipecat open STT benchmark it posts 6.99% word error rate and a 15.31% entity error rate on real agent conversations. That entity number is the one to watch for this use case, because a scam transcript is mostly entities: dollar amounts, account numbers, callback numbers, and names.
The EndOfTurn && TurnIsFormatted guard is what keeps your buffer clean. Turns arrive many times as the model revises them; you only want the last, punctuated version. Without the formatted check you'll append the same sentence three times in three states of completeness.
Terminate, then wait for the server to say it's done. Closing the socket the instant the call ends cuts off the final turn, and a fixed time.Sleep is a coin flip — the model may still be deciding whether the last utterance was a complete turn. Close flushes the buffer, sends Terminate, and then blocks until the read loop sees the server's Termination event (or five seconds pass). In testing, a 200ms sleep dropped the closing sentence of the call roughly half the time; waiting on the event has not dropped one.
Step 4: The server and the transcript buffer
Now the piece that wires Twilio to the transcriber and holds the call's text. Create main.go:
The shape here is one goroutine per call, and that's intentional. Each call gets its own transcriber, its own buffer, and its own context; when the Twilio socket dies, everything downstream unwinds. No shared map, no cleanup sweep, no leak when a call drops mid-sentence.
The media case is the hot path and it does almost nothing — decode base64, hand the bytes to the transcriber, which batches them. That's the whole point of matching Twilio's encoding instead of converting it. At 50 frames a second per call, anything expensive here becomes your scaling ceiling.
Step 5: Classify the call with the LLM Gateway
Now the analysis. Create scam.go:
Three things worth calling out.
The model is not pinned. LLM_GATEWAY_MODEL comes from the environment, and AnalyzeCall refuses to run without it. That's on purpose: the gateway's whole value is that you can move between OpenAI, Anthropic, and Google models without rewriting your integration, and a model id baked into source code is the thing that stops you doing that six months from now. Set it in config, evaluate a couple of options against your own calls, and change it when something better ships. The LLM Gateway post covers the selection story in more depth.
Both endpoints take the raw API key in the Authorization header. No Bearer prefix is needed anywhere in this tutorial — streaming and the LLM Gateway both accept the key on its own. (The gateway tolerates a Bearer prefix too, but the documented form is the bare key, so use that consistently and you have one less thing to remember.)
temperature: 0, and response_format that degrades gracefully. You want the same transcript to produce the same verdict every time. Fraud classification that wobbles between runs is worse than no classification, because someone will build an alert threshold on top of it.
JSON mode is the other half of that, and it's where the "don't pin the model" principle bites back: not every model on the gateway implements response_format. Send it to one that doesn't and the request fails outright:
If you hardcode the field, swapping to a cheaper model breaks the classifier — which defeats the point of routing through a gateway at all. So AnalyzeCall asks for JSON mode, watches for a 400 naming that field, and retries once without it. You get strict JSON where it's available and a working classifier everywhere else.
The fence-stripping before unmarshal. Even with json_object requested, some models wrap output in a markdown fence — and on the fallback path, where nothing is enforcing the shape, it's the only thing standing between you and a parse error. Three TrimPrefix/TrimSuffix calls are cheaper than an incident at 2am.
Step 6: Test it
Build and run:
With ngrok running in another terminal, call your Twilio number and read a scripted scam out loud. Something like:
"Hello, this is Michael from the fraud department at your bank. We've detected three unauthorized charges on your account ending in 4417, totaling twelve hundred dollars. To secure the account right now I need you to confirm the six-digit code we just texted you. Please stay on the line while we do this — do not hang up and do not discuss this with anyone, it's an active investigation."
Watch the server log. Partial turns stream by as you speak, a final turn lands when the model decides you've finished a thought, and when you hang up you'll get the verdict:
How many turns you get depends on how you speak. The model closes a turn on content completeness rather than on a silence timer, so reading a script straight through produces one long turn like the one above, while a real back-and-forth call produces many. Either way the buffer ends up with the same text — and note that the closing sentence, spoken right before the caller hung up, still made it in. That is the Terminate-then-wait in Close doing its job.
Notice that "4417" and "$1,200" survived intact through 8kHz telephony audio, and that the model formatted them — "twelve hundred dollars" came back as "$1,200" and "six-digit" as "6-digit". Digit strings are where telephony transcription usually falls over, and they're also the part of a fraud transcript that determines whether the verdict is actionable or just a vibe.
When it doesn't work
- The socket closes immediately and you see error 3007. You're sending Twilio's raw 20ms frames straight through. Audio messages must be 50–1000 ms; batch five frames before writing. This is what SendAudio above is for.
- Twilio connects but no transcript appears. Almost always the audio format. Confirm encoding=pcm_mulaw and sample_rate=8000 in the streaming query string, and check the start event's mediaFormat block in your logs to see what Twilio actually sent.
- streaming handshake failed (401). The key itself is wrong or unset. Both the streaming dial and the LLM Gateway call take the raw API key in the Authorization header.
- The same sentence appears three times in the buffer. You're appending on partials. The EndOfTurn && TurnIsFormatted guard is what filters those out.
- analysis failed: llm gateway returned 400 ... does not support response_format. The model in LLM_GATEWAY_MODEL doesn't implement JSON mode. The code above already retries without the field; if you stripped that fallback out, put it back or pick a model that supports it.
- Twilio's webhook 502s. ngrok restarted and the hostname changed. Update the number's voice URL and PUBLIC_HOST.
- The last sentence is missing. You closed the socket too fast, or you dropped the tail of the buffer. Close() flushes the remaining audio, sends Terminate, and then waits for the server's Termination event rather than sleeping a fixed interval. If you replace that wait with a sleep, you will lose the final turn intermittently.
Can speech-to-text APIs detect profanity or compliance issues?
Yes, and it's worth separating the two, because they're different mechanisms.
Profanity is handled at the transcription layer. Async transcription has a profanity filter that masks flagged words in the output, and content moderation in the Speech Understanding API classifies passages by sensitive topic with a confidence score. That's deterministic, cheap, and runs without an LLM.
Compliance is a judgment call, and that's what the pattern in this tutorial is for. Whether an agent read the required disclosure, whether a caller was pressured, whether a mini-Miranda was delivered before a collections conversation — none of those are keyword matches. They're questions about what was said and what it meant, which means transcript plus a model. The scam prompt above is one instance of that shape; a compliance checklist is another, and the only thing that changes is the system prompt and the JSON schema you ask for.
If you want that as a managed layer rather than a prompt you maintain, Voice AI guardrails covers the built-in version — moderation, policy enforcement, and cost controls without writing the classifier yourself.
The complete project
Three files, one dependency:
go.mod:
Run it:
The three source files above are the complete program — nothing is elided, and nothing else is imported beyond the standard library and gorilla/websocket.
Where to take it next
The version you just built classifies once, at the end. Four changes turn it into something you'd actually run.
Classify mid-call. Your OnFinalTurn callback already sees every completed turn. Run AnalyzeCall on the buffer every N turns, or on a sliding window, and you can warn a user or alert a supervisor while the scam is still in progress. The cost of that is one gateway call per evaluation, so window it rather than firing on every turn.
Add speaker labels. With two parties on a line, knowing who said "read me the code" versus who said it back changes the verdict completely. Streaming speaker diarization labels speakers live and re-clusters at the end of the stream, and the labels feed straight into the prompt.
Persist the transcript. Right now it lives in a struct and dies with the process. A row per call — SID, transcript, verdict, signals — is what turns this from a demo into something with an audit trail, which is the actual requirement in any contact center deployment.
Respond on the call. Once you're reasoning over live transcript, you're one TTS call away from a voice agent that intervenes rather than just logs. That's a different architecture and the Voice Agent API collapses it into a single WebSocket, but the transcription half is exactly what you've already written.
The part worth keeping from this build isn't the scam prompt. It's the shape: Twilio forks the audio, a WebSocket turns it into text, and a model turns the text into a decision your application can act on. Streaming transcription is billed on session duration at $0.45/hr base, which means the cost of watching a call is roughly the cost of the call being open — and the cost of not watching it is whatever the fraud was worth.
Frequently asked questions
How does Twilio real-time transcription work?
Twilio Media Streams forks a live copy of call audio to a WebSocket you control, sending base64-encoded mu-law frames at 8000 Hz roughly every 20 milliseconds. Your service decodes those frames, batches them into messages of at least 50 milliseconds, and forwards them to a speech-to-text API over a second WebSocket, which returns partial transcripts as the caller speaks and final transcripts when a turn completes. Nothing is stored in between — the transcription happens while the call is still connected.
What audio format does Twilio Media Streams send?
PCMU, also called G.711 mu-law, at 8000 Hz, mono, base64-encoded in 20ms frames of 160 bytes. Configure your streaming session with encoding=pcm_mulaw and sample_rate=8000 so you can forward the decoded bytes directly with no transcoding. Resampling to 16kHz or converting to PCM16 adds latency and CPU for no accuracy gain. Do batch the frames, though — the streaming API requires 50–1000 ms of audio per message, so a single 20ms Twilio frame is rejected on its own.
Can I transcribe Twilio calls in real time with Go?
Yes. Go needs no SDK for either side of this — a WebSocket client such as gorilla/websocket handles both the inbound Twilio media stream and the outbound connection to wss://streaming.assemblyai.com/v3/ws. The complete working service in this tutorial is about 250 lines across three files with one third-party dependency.
How much does real-time call transcription cost?
Streaming transcription with Universal-3.5 Pro Realtime is $0.45/hr base, billed per second on session duration — the time your WebSocket is open, not the duration of speech within it. LLM Gateway tokens for the classification step are priced separately. Close the streaming socket when the call ends rather than leaving it idle, since an open session bills whether audio is flowing or not.
Can a speech-to-text API flag profanity or compliance problems on a call?
Profanity, yes, directly — a profanity filter masks flagged words and content moderation classifies passages by sensitive topic with confidence scores, no LLM required. Compliance is a judgment question rather than a keyword match, so it needs the transcript plus a model reasoning over it, which is the pattern this tutorial builds. Swap the fraud system prompt for a compliance checklist and the rest of the pipeline is unchanged.
Should I use Start Stream or Connect Stream in my TwiML?
Use <Start><Stream> when you want to observe the call while it continues normally — the TwiML document keeps executing after the stream opens, so you can still dial, queue, or play audio. Use <Connect><Stream> when your WebSocket is taking over the call entirely, which is the right choice for a voice agent but the wrong one for passive monitoring.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.


