Insights & Use Cases
September 29, 2026

Detect scam calls using Go with the LLM Gateway and Twilio

Build a Go service that transcribes Twilio calls in real time and flags likely scam attempts with the AssemblyAI LLM Gateway. Full working code.

Marcus Olsson
, 
Senior Developer Educator
Reviewed by
Ryan O'Connor
, 
Senior Developer Educator
Table of contents

Twilio will happily give you a live copy of every call your number handles. That's what Media Streams does: it forks the call audio to a WebSocket you control, in real time, while the call is still going. What it won't give you is words — you get raw mulaw frames at 8kHz and it's your problem from there.

So the interesting question isn't "can I get the audio." It's what you do with it inside the two seconds before the caller finishes their sentence.

This tutorial builds the whole pipeline in Go: a Twilio number that forks its audio to your service, real-time transcription over a WebSocket, a running transcript held in memory, and an LLM pass at the end of the call that reads the transcript and decides whether it looks like a scam. Roughly 250 lines of Go, no framework, no SDK required. Everything here is complete and runnable.

Scam classification is the example because it exercises the hard parts — you need accurate entity capture (account numbers, dollar amounts, callback numbers), you need it live, and you need a model reasoning over the result. Swap the prompt and the same service does compliance monitoring, QA scoring, or intent routing.

What you'll build

Caller ──▶ Twilio number ──▶ TwiML 
                                     │
                                     ▼
                     Your Go server (WebSocket, /stream)
                                     │
                     base64 mulaw ──▶│──▶ AssemblyAI Streaming
                                     │    wss://streaming.assemblyai.com/v3/ws
                                     │◀── Turn events (partial + final)
                                     ▼
                          Transcript buffer (per call SID)
                                     │
                          call ends  ▼
                             AssemblyAI LLM Gateway
                                     │
                                     ▼
                     { verdict, confidence, signals[], summary }

Two WebSockets, one HTTP server. Twilio talks to you; you talk to AssemblyAI. The Go process sits in the middle holding state for the duration of the call.

Prerequisites

  • Go 1.21 or newer.
  • A Twilio account and a voice-capable phone number. Trial accounts work.
  • An AssemblyAI API key. Sign up free — the key covers both streaming transcription and the LLM Gateway.
  • ngrok, or any tunnel that gives you a public HTTPS/WSS URL. Twilio has to reach your laptop.

One dependency:

mkdir scamwatch && cd scamwatch
go mod init scamwatch
go get github.com/gorilla/websocket

Why real-time and not after the fact

You could record the call, transcribe it when it ends, and classify the recording. It's simpler and it's the wrong shape for this problem.

A scam call is a live event. The value of knowing "this caller is walking your user through a gift card purchase" collapses to nearly zero once the call is over and the gift cards are bought. Everything useful — warning the user mid-call, flagging the session for a supervisor, cutting the call off, dropping a compliance marker into the CRM — depends on having the text while the audio is still arriving.

The architecture also composes. Once transcript text is flowing through your process turn by turn, you can act on any turn, not just the last one. This tutorial classifies at the end of the call because that's the simplest version to reason about, but the buffer it builds is exactly what you'd tap for mid-call intervention, and the same pipeline is the front half of any voice agent you'd build on the same number.

Step 1: Get a public URL with ngrok

Twilio needs to reach your machine from the public internet, both for the TwiML webhook and for the media WebSocket. Start the tunnel first so you have the hostname before you configure anything.

ngrok http 8080

You'll get a forwarding line like:

Forwarding   https://a1b2-73-15-201-4.ngrok-free.app -> http://localhost:8080

Two URLs come out of that hostname:

Note the scheme change — https for the webhook, wss for the stream, same host. Keep this terminal open; a free ngrok hostname changes every restart and you'll have to update Twilio when it does.

Export it for the server to build TwiML with:

export PUBLIC_HOST="a1b2-73-15-201-4.ngrok-free.app"
export ASSEMBLYAI_API_KEY="your_key_here"
export LLM_GATEWAY_MODEL="the-model-id-you-picked"

That last variable is deliberately not hardcoded. The LLM Gateway fronts models from OpenAI, Anthropic, Google, and others behind one API and one bill, and the right model for a classification task like this changes as new ones ship. Pick one from the gateway's model list and set it here rather than pinning a model id into your source.

Step 2: Point a Twilio number at your server

In the Twilio Console, open Phone Numbers → Manage → Active numbers and click your number. Under Voice Configuration, set:

Save it. From now on, every inbound call to that number makes Twilio POST to /voice and execute whatever TwiML you return.

The TwiML we'll return uses <Start><Stream>, which forks a copy of the inbound audio to your WebSocket and then keeps going with the rest of the document. That's the distinction that matters: <Connect><Stream> hands the call over to your socket entirely, while <Start><Stream> taps it and lets the call proceed. We want the tap.



  This call may be monitored for quality and safety.
  
    
      
    
  
  

The <Pause> keeps the call alive so there's something to listen to. In a real deployment this is where you'd <Dial> the agent, the queue, or the voice agent that actually handles the call.

What Twilio sends you

Once the stream opens, Twilio sends JSON text frames over the WebSocket. Four event types matter:

EventWhen it firesWhat you need from it
connectedSocket handshake completedNothing — acknowledge and move on
startMedia is about to flowstart.callSid and the media format block
mediaEvery 20ms of audiomedia.payload — base64 mulaw, 8000 Hz, mono
stopCall ended or stream torn downYour cue to finalize and classify

The audio format is the part people get wrong. Twilio Media Streams are PCMU (G.711 mu-law) at 8000 Hz, mono, base64-encoded, in 20ms frames — 160 bytes of mulaw per frame. Don't resample it, don't convert it to PCM16, don't assume 16kHz. Decode the base64 and forward the bytes exactly as they arrived, then tell AssemblyAI that's what's coming.

One thing you cannot skip: batch the frames

Twilio emits a media frame every 20 milliseconds. The AssemblyAI streaming API requires between 50 ms and 1000 ms of audio per WebSocket message. Forward Twilio's frames one-for-one and the very first one is rejected:

{"type":"Error","error_code":3007,
 "error":"Input Duration Error: Input Duration Violation: 20.0 ms. Expected between 50 and 1000 ms"}

The socket then closes and you never receive a single transcript. This is the most common way this integration fails, and it fails quietly in the sense that your Twilio side looks perfectly healthy — audio is flowing, frames are arriving, and the only sign of trouble is one error frame before the connection drops.

The fix is a small buffer. Accumulate five Twilio frames (5 × 20 ms = 100 ms, 800 bytes of mu-law) and send that. 100 ms sits comfortably inside the window, adds negligible latency next to the model's own turn detection, and cuts your WebSocket write rate by 5×. The transcriber below does this internally so the rest of your code doesn't have to think about it.

Step 3: The real-time transcriber

Here's the transcriber. It opens a WebSocket to wss://streaming.assemblyai.com/v3/ws, declares the Twilio audio format in the query string, batches incoming frames to a legal message size, and dispatches turn events to callbacks.

Create transcriber.go:

package main

import (
	"context"
	"encoding/json"
	"fmt"
	"log"
	"net/http"
	"net/url"
	"sync"
	"time"

	"github.com/gorilla/websocket"
)

const streamingEndpoint = "wss://streaming.assemblyai.com/v3/ws"

// The streaming API accepts 50-1000 ms of audio per message. Twilio sends
// 20 ms frames (160 bytes of 8 kHz mu-law), so we batch five of them into a
// 100 ms message. Anything under 50 ms is rejected with error 3007 and the
// session is closed.
const (
	minSendBytes  = 400 // 50 ms  — the hard floor
	sendChunkSize = 800 // 100 ms — what we actually aim for
)

// TurnEvent is one message from the AssemblyAI streaming API. Partial turns
// arrive continuously with EndOfTurn false; the same turn arrives again with
// EndOfTurn true once the model decides the speaker has finished.
type TurnEvent struct {
	Type            string `json:"type"`
	Transcript      string `json:"transcript"`
	EndOfTurn       bool   `json:"end_of_turn"`
	TurnOrder       int    `json:"turn_order"`
	TurnIsFormatted bool   `json:"turn_is_formatted"`
	ErrorCode       int    `json:"error_code"`
	Error           string `json:"error"`
}

// RealTimeTranscriber wraps a single streaming session.
type RealTimeTranscriber struct {
	conn     *websocket.Conn
	writeMu  sync.Mutex
	pending  []byte
	closeOne sync.Once
	finished chan struct{} // closed when the read loop exits

	// OnFinalTurn fires once per completed, formatted turn.
	OnFinalTurn func(turnOrder int, text string)
	// OnPartialTurn fires on every in-progress update. Useful for live UI.
	OnPartialTurn func(text string)
}

// NewRealTimeTranscriber dials the streaming endpoint configured for Twilio
// Media Streams audio: mu-law (PCMU) at 8000 Hz, mono.
func NewRealTimeTranscriber(ctx context.Context, apiKey string) (*RealTimeTranscriber, error) {
	if apiKey == "" {
		return nil, fmt.Errorf("ASSEMBLYAI_API_KEY is not set")
	}

	q := url.Values{}
	q.Set("speech_model", "universal-3-5-pro")
	q.Set("encoding", "pcm_mulaw")
	q.Set("sample_rate", "8000")
	q.Set("format_turns", "true")

	header := http.Header{}
	header.Set("Authorization", apiKey)

	dialer := websocket.Dialer{HandshakeTimeout: 15 * time.Second}
	conn, resp, err := dialer.DialContext(ctx, streamingEndpoint+"?"+q.Encode(), header)
	if err != nil {
		if resp != nil {
			return nil, fmt.Errorf("streaming handshake failed (%s): %w", resp.Status, err)
		}
		return nil, fmt.Errorf("streaming handshake failed: %w", err)
	}

	return &RealTimeTranscriber{conn: conn, finished: make(chan struct{})}, nil
}

// SendAudio buffers one 20 ms Twilio frame and writes a batched message once
// enough audio has accumulated. Safe for concurrent use.
func (t *RealTimeTranscriber) SendAudio(frame []byte) error {
	t.writeMu.Lock()
	defer t.writeMu.Unlock()

	t.pending = append(t.pending, frame...)
	if len(t.pending) < sendChunkSize {
		return nil
	}

	chunk := t.pending
	t.pending = nil
	return t.conn.WriteMessage(websocket.BinaryMessage, chunk)
}

// Flush writes whatever is left in the buffer, padding with mu-law silence so
// the final message still clears the 50 ms floor.
func (t *RealTimeTranscriber) Flush() error {
	t.writeMu.Lock()
	defer t.writeMu.Unlock()

	if len(t.pending) == 0 {
		return nil
	}

	chunk := t.pending
	t.pending = nil
	for len(chunk) < minSendBytes {
		chunk = append(chunk, 0xFF) // 0xFF is silence in mu-law
	}
	return t.conn.WriteMessage(websocket.BinaryMessage, chunk)
}

// Listen blocks reading turn events until the socket closes or ctx is done.
func (t *RealTimeTranscriber) Listen(ctx context.Context) {
	defer close(t.finished)

	go func() {
		<-ctx.Done()
		t.Close()
	}()

	for {
		_, raw, err := t.conn.ReadMessage()
		if err != nil {
			if !websocket.IsCloseError(err, websocket.CloseNormalClosure) &&
				ctx.Err() == nil {
				log.Printf("transcriber: read error: %v", err)
			}
			return
		}

		var ev TurnEvent
		if err := json.Unmarshal(raw, &ev); err != nil {
			log.Printf("transcriber: bad frame: %v", err)
			continue
		}

		switch {
		case ev.Error != "":
			log.Printf("transcriber: api error %d: %s", ev.ErrorCode, ev.Error)
		case ev.Type == "Termination":
			// The server has flushed everything it owes us.
			return
		case ev.Type == "Turn" && ev.Transcript != "":
			if ev.EndOfTurn && ev.TurnIsFormatted {
				if t.OnFinalTurn != nil {
					t.OnFinalTurn(ev.TurnOrder, ev.Transcript)
				}
			} else if t.OnPartialTurn != nil {
				t.OnPartialTurn(ev.Transcript)
			}
		}
	}
}

// Close flushes, sends the terminate message and shuts the socket down once.
func (t *RealTimeTranscriber) Close() {
	t.closeOne.Do(func() {
		_ = t.Flush()

		t.writeMu.Lock()
		_ = t.conn.WriteJSON(map[string]string{"type": "Terminate"})
		t.writeMu.Unlock()

		// Wait for the server to emit any remaining turns and terminate the
		// session. A fixed sleep here is a coin flip -- the model may still be
		// deciding whether the last utterance is a complete turn, and if it is
		// still thinking when we hang up, that turn is lost.
		select {
		case <-t.finished:
		case <-time.After(5 * time.Second):
			log.Printf("transcriber: timed out waiting for final turns")
		}

		_ = t.conn.Close()
	})
}

A few notes on the choices in there.

encoding=pcm_mulaw and sample_rate=8000 are doing real work. The streaming API accepts mulaw natively, which means zero transcoding in your hot path — the bytes Twilio hands you are the bytes you write to the socket. Get either parameter wrong and you won't get an error; you'll get a transcript that reads like static.

The batching in SendAudio is load-bearing, not an optimization. Without it the session dies on the first frame with error 3007. Keeping the buffer inside the transcriber means the Twilio handler stays a straight decode-and-forward loop, and it means Flush on close doesn't lose the tail of the call.

speech_model=universal-3-5-pro selects Universal-3.5 Pro Realtime, the current streaming flagship. Telephony audio at 8kHz is the hardest input a speech model gets — it's band-limited, often compressed twice, and frequently noisy — which is exactly the case where model choice shows up in the output. On the Pipecat open STT benchmark it posts 6.99% word error rate and a 15.31% entity error rate on real agent conversations. That entity number is the one to watch for this use case, because a scam transcript is mostly entities: dollar amounts, account numbers, callback numbers, and names.

The EndOfTurn && TurnIsFormatted guard is what keeps your buffer clean. Turns arrive many times as the model revises them; you only want the last, punctuated version. Without the formatted check you'll append the same sentence three times in three states of completeness.

Terminate, then wait for the server to say it's done. Closing the socket the instant the call ends cuts off the final turn, and a fixed time.Sleep is a coin flip — the model may still be deciding whether the last utterance was a complete turn. Close flushes the buffer, sends Terminate, and then blocks until the read loop sees the server's Termination event (or five seconds pass). In testing, a 200ms sleep dropped the closing sentence of the call roughly half the time; waiting on the event has not dropped one.

Transcribe Phone Calls In Real Time

One WebSocket, native mu-law support, and accuracy that holds up on 8kHz telephony audio. Get a free API key and stream your first call in minutes.

Sign up free

Step 4: The server and the transcript buffer

Now the piece that wires Twilio to the transcriber and holds the call's text. Create main.go:

package main

import (
	"context"
	"encoding/base64"
	"encoding/json"
	"fmt"
	"log"
	"net/http"
	"os"
	"strings"
	"sync"

	"github.com/gorilla/websocket"
)

// twilioFrame is the envelope Twilio Media Streams sends over the socket.
type twilioFrame struct {
	Event string `json:"event"`
	Start struct {
		CallSid     string `json:"callSid"`
		StreamSid   string `json:"streamSid"`
		MediaFormat struct {
			Encoding   string `json:"encoding"`
			SampleRate int    `json:"sampleRate"`
			Channels   int    `json:"channels"`
		} `json:"mediaFormat"`
	} `json:"start"`
	Media struct {
		Payload string `json:"payload"`
	} `json:"media"`
}

// callTranscript accumulates finalized turns for one call.
type callTranscript struct {
	mu      sync.Mutex
	callSid string
	turns   []string
}

func (c *callTranscript) SetCallSid(sid string) {
	c.mu.Lock()
	defer c.mu.Unlock()
	c.callSid = sid
}

func (c *callTranscript) CallSid() string {
	c.mu.Lock()
	defer c.mu.Unlock()
	return c.callSid
}

func (c *callTranscript) Append(text string) {
	c.mu.Lock()
	defer c.mu.Unlock()
	c.turns = append(c.turns, text)
}

func (c *callTranscript) String() string {
	c.mu.Lock()
	defer c.mu.Unlock()
	return strings.Join(c.turns, " ")
}

var upgrader = websocket.Upgrader{
	ReadBufferSize:  4096,
	WriteBufferSize: 4096,
	CheckOrigin:     func(r *http.Request) bool { return true },
}

func main() {
	apiKey := os.Getenv("ASSEMBLYAI_API_KEY")
	if apiKey == "" {
		log.Fatal("set ASSEMBLYAI_API_KEY")
	}
	publicHost := os.Getenv("PUBLIC_HOST")
	if publicHost == "" {
		log.Fatal("set PUBLIC_HOST to your ngrok hostname, e.g. abc123.ngrok-free.app")
	}

	http.HandleFunc("/voice", func(w http.ResponseWriter, r *http.Request) {
		twiml := fmt.Sprintf(`

  This call may be monitored for quality and safety.
  
    
      
    
  
  
`, publicHost)

		w.Header().Set("Content-Type", "application/xml")
		_, _ = w.Write([]byte(twiml))
	})

	http.HandleFunc("/stream", func(w http.ResponseWriter, r *http.Request) {
		conn, err := upgrader.Upgrade(w, r, nil)
		if err != nil {
			log.Printf("upgrade failed: %v", err)
			return
		}
		handleCall(r.Context(), conn, apiKey)
	})

	log.Println("listening on :8080")
	log.Fatal(http.ListenAndServe(":8080", nil))
}

// handleCall owns one Twilio media stream from connect to stop.
func handleCall(parent context.Context, twilioConn *websocket.Conn, apiKey string) {
	defer twilioConn.Close()

	ctx, cancel := context.WithCancel(parent)
	defer cancel()

	transcriber, err := NewRealTimeTranscriber(ctx, apiKey)
	if err != nil {
		log.Printf("could not start transcriber: %v", err)
		return
	}
	defer transcriber.Close()

	record := &callTranscript{}

	transcriber.OnFinalTurn = func(order int, text string) {
		log.Printf("[%s] turn %d: %s", record.CallSid(), order, text)
		record.Append(text)
	}
	transcriber.OnPartialTurn = func(text string) {
		// Swap this for a websocket push if you want live captions in a UI.
		log.Printf("[%s] ... %s", record.CallSid(), text)
	}

	go transcriber.Listen(ctx)

	for done := false; !done; {
		_, raw, err := twilioConn.ReadMessage()
		if err != nil {
			log.Printf("twilio socket closed: %v", err)
			break
		}

		var frame twilioFrame
		if err := json.Unmarshal(raw, &frame); err != nil {
			continue
		}

		switch frame.Event {
		case "connected":
			log.Println("twilio: connected")

		case "start":
			record.SetCallSid(frame.Start.CallSid)
			log.Printf("twilio: stream started for call %s (%s, %d Hz, %d ch)",
				frame.Start.CallSid,
				frame.Start.MediaFormat.Encoding,
				frame.Start.MediaFormat.SampleRate,
				frame.Start.MediaFormat.Channels,
			)

		case "media":
			audio, err := base64.StdEncoding.DecodeString(frame.Media.Payload)
			if err != nil {
				continue
			}
			if err := transcriber.SendAudio(audio); err != nil {
				log.Printf("forward failed: %v", err)
			}

		case "stop":
			log.Printf("twilio: stream stopped for call %s", record.CallSid())
			done = true
		}
	}

	cancel()
	transcriber.Close()

	full := record.String()
	if strings.TrimSpace(full) == "" {
		log.Printf("[%s] no speech captured, skipping analysis", record.CallSid())
		return
	}

	result, err := AnalyzeCall(context.Background(), apiKey, full)
	if err != nil {
		log.Printf("[%s] analysis failed: %v", record.CallSid(), err)
		return
	}

	out, _ := json.MarshalIndent(result, "", "  ")
	log.Printf("[%s] verdict:\n%s", record.CallSid(), out)
}

The shape here is one goroutine per call, and that's intentional. Each call gets its own transcriber, its own buffer, and its own context; when the Twilio socket dies, everything downstream unwinds. No shared map, no cleanup sweep, no leak when a call drops mid-sentence.

The media case is the hot path and it does almost nothing — decode base64, hand the bytes to the transcriber, which batches them. That's the whole point of matching Twilio's encoding instead of converting it. At 50 frames a second per call, anything expensive here becomes your scaling ceiling.

Step 5: Classify the call with the LLM Gateway

Now the analysis. Create scam.go:

package main

import (
	"bytes"
	"context"
	"encoding/json"
	"errors"
	"fmt"
	"io"
	"net/http"
	"os"
	"strings"
	"time"
)

const llmGatewayURL = "https://llm-gateway.assemblyai.com/v1/chat/completions"

// CallAnalysis is the structured verdict we ask the model to produce.
type CallAnalysis struct {
	Verdict      string   `json:"verdict"`       // "scam" | "suspicious" | "legitimate"
	Confidence   float64  `json:"confidence"`    // 0.0 - 1.0
	Signals      []string `json:"signals"`       // observed tactics
	EntitiesSeen []string `json:"entities_seen"` // amounts, accounts, callbacks
	Summary      string   `json:"summary"`       // one or two sentences
}

type chatMessage struct {
	Role    string `json:"role"`
	Content string `json:"content"`
}

type responseFormat struct {
	Type string `json:"type"`
}

type chatRequest struct {
	Model    string        `json:"model"`
	Messages []chatMessage `json:"messages"`
	// Temperature is a pointer so an explicit 0 is still sent.
	Temperature *float64 `json:"temperature,omitempty"`
	// Omitted entirely when nil — not every gateway model accepts it.
	ResponseFormat *responseFormat `json:"response_format,omitempty"`
}

type chatResponse struct {
	Choices []struct {
		Message chatMessage `json:"message"`
	} `json:"choices"`
	Error *struct {
		Message string `json:"message"`
	} `json:"error"`
}

const systemPrompt = `You are a fraud analyst reviewing a transcript of a single
phone call. Classify the call and explain your reasoning from the transcript only.

Known scam patterns to look for:
- Manufactured urgency ("your account will be closed in the next hour")
- Impersonation of a bank, tax authority, utility, or law enforcement
- Requests for gift cards, wire transfers, crypto, or remote desktop access
- Requests for one-time passcodes, full card numbers, or account credentials
- Instructions to keep the call secret or to stay on the line while acting
- A callback number that differs from the organization being claimed

Return ONLY a JSON object with these keys:
  verdict        one of "scam", "suspicious", "legitimate"
  confidence     number between 0 and 1
  signals        array of short strings naming the patterns you observed
  entities_seen  array of sensitive values mentioned (amounts, account refs,
                 callback numbers), redacted to their last 4 characters
  summary        one or two sentences describing what the caller wanted

If the transcript is too short or too garbled to judge, return verdict
"legitimate" with a confidence below 0.3 and say so in the summary.`

// errNoResponseFormat means the chosen model rejected the response_format
// field. Plenty of models on the gateway don't implement JSON mode.
var errNoResponseFormat = errors.New("model does not support response_format")

// postChat performs one gateway call and returns the assistant's content.
func postChat(ctx context.Context, apiKey string, reqBody chatRequest) (string, error) {
	payload, err := json.Marshal(reqBody)
	if err != nil {
		return "", err
	}

	req, err := http.NewRequestWithContext(ctx, http.MethodPost, llmGatewayURL, bytes.NewReader(payload))
	if err != nil {
		return "", err
	}
	req.Header.Set("Authorization", apiKey)
	req.Header.Set("Content-Type", "application/json")

	client := &http.Client{Timeout: 60 * time.Second}
	resp, err := client.Do(req)
	if err != nil {
		return "", err
	}
	defer resp.Body.Close()

	body, err := io.ReadAll(resp.Body)
	if err != nil {
		return "", err
	}
	if resp.StatusCode != http.StatusOK {
		if resp.StatusCode == http.StatusBadRequest &&
			bytes.Contains(body, []byte("response_format")) {
			return "", errNoResponseFormat
		}
		return "", fmt.Errorf("llm gateway returned %d: %s", resp.StatusCode, string(body))
	}

	var parsed chatResponse
	if err := json.Unmarshal(body, &parsed); err != nil {
		return "", fmt.Errorf("could not parse gateway response: %w", err)
	}
	if parsed.Error != nil {
		return "", fmt.Errorf("llm gateway error: %s", parsed.Error.Message)
	}
	if len(parsed.Choices) == 0 {
		return "", fmt.Errorf("llm gateway returned no choices")
	}

	return parsed.Choices[0].Message.Content, nil
}

// AnalyzeCall sends the finished transcript through the LLM Gateway.
func AnalyzeCall(ctx context.Context, apiKey, transcript string) (*CallAnalysis, error) {
	model := os.Getenv("LLM_GATEWAY_MODEL")
	if model == "" {
		return nil, fmt.Errorf("set LLM_GATEWAY_MODEL to a model id the LLM Gateway exposes")
	}

	zero := 0.0
	reqBody := chatRequest{
		Model:       model,
		Temperature: &zero,
		Messages: []chatMessage{
			{Role: "system", Content: systemPrompt},
			{Role: "user", Content: "Call transcript:\n\n" + transcript},
		},
		ResponseFormat: &responseFormat{Type: "json_object"},
	}

	// Ask for JSON mode first, and drop it if this model doesn't implement it.
	// The system prompt already demands a bare JSON object, and the fence
	// stripping below handles the rest, so the fallback costs us nothing but
	// strictness.
	content, err := postChat(ctx, apiKey, reqBody)
	if errors.Is(err, errNoResponseFormat) {
		reqBody.ResponseFormat = nil
		content, err = postChat(ctx, apiKey, reqBody)
	}
	if err != nil {
		return nil, err
	}

	content = strings.TrimSpace(content)
	content = strings.TrimPrefix(content, "```json")
	content = strings.TrimPrefix(content, "```")
	content = strings.TrimSuffix(content, "```")

	var analysis CallAnalysis
	if err := json.Unmarshal([]byte(strings.TrimSpace(content)), &analysis); err != nil {
		return nil, fmt.Errorf("model did not return valid JSON: %w (raw: %s)", err, content)
	}

	return &analysis, nil
}

Three things worth calling out.

The model is not pinned. LLM_GATEWAY_MODEL comes from the environment, and AnalyzeCall refuses to run without it. That's on purpose: the gateway's whole value is that you can move between OpenAI, Anthropic, and Google models without rewriting your integration, and a model id baked into source code is the thing that stops you doing that six months from now. Set it in config, evaluate a couple of options against your own calls, and change it when something better ships. The LLM Gateway post covers the selection story in more depth.

Both endpoints take the raw API key in the Authorization header. No Bearer prefix is needed anywhere in this tutorial — streaming and the LLM Gateway both accept the key on its own. (The gateway tolerates a Bearer prefix too, but the documented form is the bare key, so use that consistently and you have one less thing to remember.)

temperature: 0, and response_format that degrades gracefully. You want the same transcript to produce the same verdict every time. Fraud classification that wobbles between runs is worse than no classification, because someone will build an alert threshold on top of it.

JSON mode is the other half of that, and it's where the "don't pin the model" principle bites back: not every model on the gateway implements response_format. Send it to one that doesn't and the request fails outright:

{"metadata":{"errors":["model qwen3.5-4b-32k-fast does not support response_format"]},
 "message":"invalid request body","code":400}

If you hardcode the field, swapping to a cheaper model breaks the classifier — which defeats the point of routing through a gateway at all. So AnalyzeCall asks for JSON mode, watches for a 400 naming that field, and retries once without it. You get strict JSON where it's available and a working classifier everywhere else.

The fence-stripping before unmarshal. Even with json_object requested, some models wrap output in a markdown fence — and on the fallback path, where nothing is enforcing the shape, it's the only thing standing between you and a parse error. Three TrimPrefix/TrimSuffix calls are cheaper than an incident at 2am.

Step 6: Test it

Build and run:

go build -o scamwatch . && ./scamwatch

With ngrok running in another terminal, call your Twilio number and read a scripted scam out loud. Something like:

"Hello, this is Michael from the fraud department at your bank. We've detected three unauthorized charges on your account ending in 4417, totaling twelve hundred dollars. To secure the account right now I need you to confirm the six-digit code we just texted you. Please stay on the line while we do this — do not hang up and do not discuss this with anyone, it's an active investigation."

Watch the server log. Partial turns stream by as you speak, a final turn lands when the model decides you've finished a thought, and when you hang up you'll get the verdict:

2026/09/29 11:04:12 twilio: connected
2026/09/29 11:04:12 twilio: stream started for call CA9f3... (audio/x-mulaw, 8000 Hz, 1 ch)
2026/09/29 11:04:13 [CA9f3...] ... Hello, this
2026/09/29 11:04:16 [CA9f3...] ... Hello, this is Michael from the fraud department at your bank.
2026/09/29 11:04:22 [CA9f3...] ... Hello, this is Michael from the fraud department at your bank. We have detected 3 unauthorized charges on your account ending in 4417,
2026/09/29 11:04:31 [CA9f3...] ... Hello, this is Michael from the fraud department at your bank. We have detected 3 unauthorized charges on your account ending in 4417, totaling $1,200. To secure the account right now. I need you to confirm the 6-digit code we just texted you.
2026/09/29 11:04:33 twilio: stream stopped for call CA9f3...
2026/09/29 11:04:34 [CA9f3...] turn 0: Hello, this is Michael from the fraud department at your bank. We have detected 3 unauthorized charges on your account ending in 4417, totaling $1,200. To secure the account right now. I need you to confirm the 6-digit code we just texted you. Please stay on the line while we do this.
2026/09/29 11:04:36 [CA9f3...] verdict:
{
  "verdict": "scam",
  "confidence": 0.95,
  "signals": [
    "Impersonation of a bank",
    "Manufactured urgency",
    "Requests for one-time passcodes",
    "Instructions to stay on the line"
  ],
  "entities_seen": [
    "4417",
    "1200",
    "6-digit code"
  ],
  "summary": "The caller impersonated a bank fraud department to create urgency and requested the victim provide a 6-digit code sent via text to secure an account."
}

How many turns you get depends on how you speak. The model closes a turn on content completeness rather than on a silence timer, so reading a script straight through produces one long turn like the one above, while a real back-and-forth call produces many. Either way the buffer ends up with the same text — and note that the closing sentence, spoken right before the caller hung up, still made it in. That is the Terminate-then-wait in Close doing its job.

Notice that "4417" and "$1,200" survived intact through 8kHz telephony audio, and that the model formatted them — "twelve hundred dollars" came back as "$1,200" and "six-digit" as "6-digit". Digit strings are where telephony transcription usually falls over, and they're also the part of a fraud transcript that determines whether the verdict is actionable or just a vibe.

When it doesn't work

  • The socket closes immediately and you see error 3007. You're sending Twilio's raw 20ms frames straight through. Audio messages must be 50–1000 ms; batch five frames before writing. This is what SendAudio above is for.
  • Twilio connects but no transcript appears. Almost always the audio format. Confirm encoding=pcm_mulaw and sample_rate=8000 in the streaming query string, and check the start event's mediaFormat block in your logs to see what Twilio actually sent.
  • streaming handshake failed (401). The key itself is wrong or unset. Both the streaming dial and the LLM Gateway call take the raw API key in the Authorization header.
  • The same sentence appears three times in the buffer. You're appending on partials. The EndOfTurn && TurnIsFormatted guard is what filters those out.
  • analysis failed: llm gateway returned 400 ... does not support response_format. The model in LLM_GATEWAY_MODEL doesn't implement JSON mode. The code above already retries without the field; if you stripped that fallback out, put it back or pick a model that supports it.
  • Twilio's webhook 502s. ngrok restarted and the hostname changed. Update the number's voice URL and PUBLIC_HOST.
  • The last sentence is missing. You closed the socket too fast, or you dropped the tail of the buffer. Close() flushes the remaining audio, sends Terminate, and then waits for the server's Termination event rather than sleeping a fixed interval. If you replace that wait with a sleep, you will lose the final turn intermittently.
Test Streaming On Your Own Calls

Run your own telephony audio through streaming transcription and see how it handles digits, names, and background noise before you write a line of code.

Try playground

Can speech-to-text APIs detect profanity or compliance issues?

Yes, and it's worth separating the two, because they're different mechanisms.

Profanity is handled at the transcription layer. Async transcription has a profanity filter that masks flagged words in the output, and content moderation in the Speech Understanding API classifies passages by sensitive topic with a confidence score. That's deterministic, cheap, and runs without an LLM.

Compliance is a judgment call, and that's what the pattern in this tutorial is for. Whether an agent read the required disclosure, whether a caller was pressured, whether a mini-Miranda was delivered before a collections conversation — none of those are keyword matches. They're questions about what was said and what it meant, which means transcript plus a model. The scam prompt above is one instance of that shape; a compliance checklist is another, and the only thing that changes is the system prompt and the JSON schema you ask for.

If you want that as a managed layer rather than a prompt you maintain, Voice AI guardrails covers the built-in version — moderation, policy enforcement, and cost controls without writing the classifier yourself.

The complete project

Three files, one dependency:

scamwatch/
├── go.mod
├── main.go          # HTTP server, TwiML, Twilio socket, transcript buffer
├── transcriber.go   # RealTimeTranscriber — AssemblyAI streaming WebSocket
└── scam.go          # AnalyzeCall — LLM Gateway classification

go.mod:

module scamwatch

go 1.21

require github.com/gorilla/websocket v1.5.3

Run it:

export ASSEMBLYAI_API_KEY="your_key"
export PUBLIC_HOST="your-subdomain.ngrok-free.app"
export LLM_GATEWAY_MODEL="the-model-id-you-picked"

go build -o scamwatch . && ./scamwatch

The three source files above are the complete program — nothing is elided, and nothing else is imported beyond the standard library and gorilla/websocket.

Where to take it next

The version you just built classifies once, at the end. Four changes turn it into something you'd actually run.

Classify mid-call. Your OnFinalTurn callback already sees every completed turn. Run AnalyzeCall on the buffer every N turns, or on a sliding window, and you can warn a user or alert a supervisor while the scam is still in progress. The cost of that is one gateway call per evaluation, so window it rather than firing on every turn.

Add speaker labels. With two parties on a line, knowing who said "read me the code" versus who said it back changes the verdict completely. Streaming speaker diarization labels speakers live and re-clusters at the end of the stream, and the labels feed straight into the prompt.

Persist the transcript. Right now it lives in a struct and dies with the process. A row per call — SID, transcript, verdict, signals — is what turns this from a demo into something with an audit trail, which is the actual requirement in any contact center deployment.

Respond on the call. Once you're reasoning over live transcript, you're one TTS call away from a voice agent that intervenes rather than just logs. That's a different architecture and the Voice Agent API collapses it into a single WebSocket, but the transcription half is exactly what you've already written.

The part worth keeping from this build isn't the scam prompt. It's the shape: Twilio forks the audio, a WebSocket turns it into text, and a model turns the text into a decision your application can act on. Streaming transcription is billed on session duration at $0.45/hr base, which means the cost of watching a call is roughly the cost of the call being open — and the cost of not watching it is whatever the fraud was worth.

Build It On Your Own Number

Get a free API key, point a Twilio number at your service, and have live transcripts flowing in an afternoon. Full API reference and SDKs included.

Sign up free

Frequently asked questions

How does Twilio real-time transcription work?

Twilio Media Streams forks a live copy of call audio to a WebSocket you control, sending base64-encoded mu-law frames at 8000 Hz roughly every 20 milliseconds. Your service decodes those frames, batches them into messages of at least 50 milliseconds, and forwards them to a speech-to-text API over a second WebSocket, which returns partial transcripts as the caller speaks and final transcripts when a turn completes. Nothing is stored in between — the transcription happens while the call is still connected.

What audio format does Twilio Media Streams send?

PCMU, also called G.711 mu-law, at 8000 Hz, mono, base64-encoded in 20ms frames of 160 bytes. Configure your streaming session with encoding=pcm_mulaw and sample_rate=8000 so you can forward the decoded bytes directly with no transcoding. Resampling to 16kHz or converting to PCM16 adds latency and CPU for no accuracy gain. Do batch the frames, though — the streaming API requires 50–1000 ms of audio per message, so a single 20ms Twilio frame is rejected on its own.

Can I transcribe Twilio calls in real time with Go?

Yes. Go needs no SDK for either side of this — a WebSocket client such as gorilla/websocket handles both the inbound Twilio media stream and the outbound connection to wss://streaming.assemblyai.com/v3/ws. The complete working service in this tutorial is about 250 lines across three files with one third-party dependency.

How much does real-time call transcription cost?

Streaming transcription with Universal-3.5 Pro Realtime is $0.45/hr base, billed per second on session duration — the time your WebSocket is open, not the duration of speech within it. LLM Gateway tokens for the classification step are priced separately. Close the streaming socket when the call ends rather than leaving it idle, since an open session bills whether audio is flowing or not.

Can a speech-to-text API flag profanity or compliance problems on a call?

Profanity, yes, directly — a profanity filter masks flagged words and content moderation classifies passages by sensitive topic with confidence scores, no LLM required. Compliance is a judgment question rather than a keyword match, so it needs the transcript plus a model reasoning over it, which is the pattern this tutorial builds. Swap the fraud system prompt for a compliance checklist and the rest of the pipeline is unchanged.

Should I use Start Stream or Connect Stream in my TwiML?

Use <Start><Stream> when you want to observe the call while it continues normally — the TwiML document keeps executing after the stream opens, so you can still dial, queue, or play audio. Use <Connect><Stream> when your WebSocket is taking over the call entirely, which is the right choice for a voice agent but the wrong one for passive monitoring.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Tutorial
LeMUR