Skip to content

Monitoring & observability

You can see what SIPhon is doing four ways: Prometheus metrics (built-in + your own), the admin API, Call Detail Records, and full SIP tracing to Homer. None of them block the call path.

Prometheus metrics

Enable the endpoint:

metrics:
  prometheus:
    listen: "0.0.0.0:9090"
    path: "/metrics"

SIPhon exports built-in gauges/counters; the ones worth alerting on:

Signal Alert when Why
siphon_memory_allocated_bytes rate(...[30m]) > 0 at flat call rate A real memory leak
siphon_pyexec_jobs_shed_total sustained rate() > 0 Handler pool saturated → SIP retransmits
siphon_pyexec_pool_size vs _pool_max pinned equal + all busy for minutes Pool fully grown and saturated
siphon_proxy_dialog_sessions grows under flat completed-call load Dialog state not draining
siphon_rtpengine_instances_up drops below your engine count An RTPEngine is unhealthy
siphon_diameter_peer_up{peer} == 0 That named peer is down. The bare siphon_diameter_peers_connected count cannot say which one
siphon_diameter_answers_total{result_code!~"2.*"} rate() > 0 A peer is reachable and refusing. siphon_diameter_request_errors_total counts transport failures only, so this reads zero there
siphon_ro_denials_total rate() > 0 Calls being refused credit by the OCS
siphon_ro_credit_teardowns_total{reason="no_teardown_hook"} increase() > 0 Credit ran out with nothing wired to enforce it — the call is still up, and unpaid
siphon_diameter_inbound_answers_total{result_code="3002"} rate() > 0 Server role only: siphon is rejecting inbound requests because no @diameter.on_request handler matched — a script gap, not a peer problem
siphon_diameter_inbound_answers_total{result_code="5012"} rate() > 0 Server role only: an @diameter.on_request handler raised or returned the wrong type. Distinct from 3002 on purpose — on Ro a 3002 reads as a credit denial and tears the call down
siphon_gateway_source_last_success_timestamp_seconds time() - <gauge> > 3 * refresh_secs, or == 0 The carriers are stale: siphon kept the last set it read and the controller has been unreadable since. 0 = never read, so a node that booted into a controller outage has no carriers from the source
siphon_gateway_source_failures_total sustained rate() > 0 Reads are failing now; the timestamp above says for how long
siphon_b2bua_inbound_calls_refused_total{reason} sustained rate() > 0 Inbound calls are being turned away at b2bua.inbound_limit. reason="concurrent" is the instance full, reason="rate" is a burst arriving faster than max_calls_per_second
siphon_b2bua_calls_active / siphon_b2bua_max_concurrent_calls ratio above your headroom, e.g. > 0.8 Capacity is running out before calls are refused. The limit gauge reads 0 when no ceiling is set, so guard the division
siphon_gateway_inbound_calls_refused_total{group,reason} sustained rate() > 0 A carrier is being turned away at its group's inbound_limit. Either it is sending more than was agreed, or the limit is set below what was
siphon_gateway_inbound_calls_active{group} / siphon_gateway_inbound_max_concurrent_calls{group} ratio above your headroom A carrier is close to its ceiling. Only groups with an inbound_limit have these series, and they are refreshed every 30 s
siphon_registrant_source_last_success_timestamp_seconds / _failures_total same pair Outbound trunk registrations, tracked separately — a node can route over stale carriers with current registrations, or the reverse

See Handler execution model for the pool internals.

Your own metrics

The metrics namespace adds counters, gauges, and histograms that appear on the same /metrics endpoint:

from siphon import metrics

calls = metrics.counter("calls_total", "Calls processed", labels=["direction", "result"])
active = metrics.gauge("calls_active", "Active calls", labels=["direction"])
setup  = metrics.histogram("call_setup_seconds", "INVITE→200 latency",
                           buckets=[0.1, 0.25, 0.5, 1, 2.5, 5])

calls.labels(direction="outbound", result="ok").inc()
active.labels(direction="outbound").inc()      # ... .dec() when it ends
setup.observe(0.342)

Admin API — health, readiness, registrations

A separate HTTP port for probes and runtime inspection:

admin:
  listen: "0.0.0.0:9091"
Endpoint Use
GET /admin/health liveness — 200 while the process is alive (survives drain)
GET /admin/ready readiness — 200, or 503 while draining (SIGTERM)
GET /admin/stats uptime + active registration count
GET /admin/registrations[/{aor}] inspect bindings
DELETE /admin/registrations/{aor} force-unregister
GET /admin/bans / DELETE /admin/bans/{ip} list / lift auto-bans (404 on both when security.failed_auth_ban is off, so an empty list always means "watching, nothing banned")
GET /admin/logs retained log lines (WARN+, plus down to admin.log_tail.retain_level); filters level, contains, call_id, paging limit + before=<seq>
GET /admin/gateways per-group dispatcher status (destinations, health, weight, priority, missed health-checks)
POST /admin/gateways/{group}/{destination}/{up\|down} mark a gateway destination up/down (drain / restore a carrier)
POST /admin/registrants/refresh re-read a registrant.backend: database / http source and reconcile now
POST /admin/gateways/refresh re-read a gateway.backend: database / http source and reconcile now
GET /admin/calls active B2BUA calls (Call-ID, state, caller, callee, B-legs)
GET /admin/metrics.json curated JSON snapshot of the live gauges + counters

Point Kubernetes liveness at /admin/health and readiness at /admin/ready so a draining pod leaves rotation cleanly — see Deployment & operations.

Bearer-token auth

The admin API can force-unregister bindings and lift bans, so gate it once it is reachable by anything but localhost:

admin:
  listen: "127.0.0.1:9091"
  auth:
    token: "${ADMIN_TOKEN}"     # keep the literal out of YAML
    protect_reads: false         # true = also require it on GET + /metrics

With a token set, the DELETE routes require Authorization: Bearer <token> (constant-time compared). Reads stay open unless protect_reads is true. Unset leaves the API open, exactly as before.

Web dashboard (experimental)

A single-page operator dashboard is baked into the binary and served same-origin on the admin listener. It's experimental — expect changes. The release Docker image compiles it in; you just enable it in config:

admin:
  listen: "127.0.0.1:9091"
  ui:
    enabled: true

A plain cargo build leaves the ui feature off, so any project embedding siphon as a library carries none of it; a binary built without --features ui logs a warning and serves nothing when enabled is set. Serving the dashboard logs an EXPERIMENTAL warning.

The views are Overview, Calls, Registrations, Gateways, Signalling, Media, Control, Security and System, reading /admin/metrics.json, /admin/registrations, /admin/calls, /admin/gateways and /admin/bans. Force-unregister, lift-ban and gateway mark-up/down go through the same bearer token — click Unlock and paste the admin.auth.token. Bind the listener internally and put it behind your own ingress auth for anything beyond a trusted network.

The dashboard is deliberately not a Grafana replacement. Rates and history belong in Prometheus; what the dashboard shows is the live, listable state Prometheus cannot hold — which calls are up, which AoR is bound where, which gateway is down, which control app owns which channel. You cannot put a Call-ID in a Prometheus label.

A subsystem this node has not configured shows as "not configured" rather than a zero, and its nav entry is greyed. That distinction matters: a permanent zero reads exactly like an idle subsystem, which is how a set of never-written metrics went unnoticed on this dashboard for a long time.

Reading the traffic counters

siphon_requests_total{method,direction} and siphon_responses_total{class,direction} count wire events, not transactions. Retransmit detection happens well downstream of the point they are taken, so a UDP INVITE retransmitted by Timer A counts once per datagram, and a non-2xx INVITE response retransmitted until ACK counts once per send. That is the right meaning for a receive/send counter, but it means requests_total is not a call count and not a transaction count — use siphon_dialogs_active (or its siphon_proxy_dialog_sessions / siphon_b2bua_calls_active halves) for concurrency, and CDRs for completed calls.

A method siphon does not recognise is counted under OTHER rather than by its own name. The method is read straight off the request line, so labelling by it would let any peer mint unbounded Prometheus series through the scrape endpoint.

siphon_connections_active{transport} has no UDP series. UDP is connectionless, so there is no connection to count; a zero there would read as "no UDP traffic", which is a different and wrong statement.

siphon_non_sip_datagrams_dropped_total{reason} is not an abuse counter and feeds nothing into failed_auth_ban. It counts payloads siphon refused to hand the SIP parser because they cannot be a SIP message: whitespace (RFC 3261 §7.5 / RFC 5626 §4.4.1, any transport), all_nul (a vendor NAT keepalive no RFC defines, sent at a registered contact every few seconds for the life of the binding), and too_short (a UDP datagram under the 14-byte grammar floor for a SIP start line). A steady all_nul rate proportional to your registration count is normal. What is worth an alert is a step change — a peer changed behaviour — or a too_short rate with no keepalive behind it, which is worth one RUST_LOG=siphon::dispatcher=trace session to see the bytes. Anything the parser does reject still warns, but only once per source per minute, with the next line carrying the count that went unlogged.

Call Detail Records

cdr:
  enabled: true
  auto_emit: true            # write one CDR per call automatically
  include_register: false    # with auto_emit, also emit a CDR per REGISTER
  backend: http              # file | http | syslog
  http:
    url: "https://collector.example.com/v1/cdr"
    auth_header: "Bearer tok123"

CDRs are written asynchronously (a bounded channel, never blocks a call) with the call's timing, parties, transport, disconnect initiator, and response code.

With auto_emit: true siphon writes one CDR per call on its own — proxy or B2BUA, no script needed — filling in timestamp_start/answer/end, duration_secs, response_code, and disconnect_initiator (caller / callee / timeout / error). Answered calls, B-leg failures, answer timeouts and caller CANCELs all produce a record. It defaults off, so it never surprises a manual-only setup.

You can also write records from a script — either instead of, or on top of, auto_emit (use it to attach billing_id / trunk / account fields). cdr.write() takes the proxy request or, from a B2BUA handler, the call:

from siphon import cdr

@proxy.on_request("INVITE")
def route(request):
    cdr.write(request, extra={"billing_id": "B-12345", "account": "ACC-789"})

@b2bua.on_answer
def answered(call, reply):
    cdr.write(call, extra={"billing_id": "B-12345"})

With auto_emit on, that does not write a second record: siphon is already tracking one for the call, and extra is merged into it, so the fields land on the same row as duration_secs, response_code and disconnect_initiator when the call ends. Call it as often as you like — later fields merge on top, last write wins per key. A record of its own is written only when there is nothing to merge into: auto_emit off, or a request the auto-emit hooks do not track (a MESSAGE, an out-of-dialog request).

So attach fields as soon as you know them (@proxy.on_request("INVITE"), @b2bua.on_answer) rather than trying to assemble a record at teardown — the teardown half is siphon's job.

A call the proxy never forwards (the script answered a 403, or dropped it) normally produces no CDR at all. Calling cdr.write() on it is how you ask for one anyway — a blocked-call record, emitted with the code the script answered.

Watch siphon_cdr_sessions (the live per-call tracking count): it returns to 0 between calls, and a steady climb under flat load means a call teardown isn't being seen.

Full SIP tracing → Homer

Stream every SIP message to a Homer / heplify-server collector over HEP — invaluable for debugging call flows:

tracing:
  hep:
    endpoint: "127.0.0.1:9060"
    version: 3
    transport: udp           # udp | tcp | tls
    agent_id: "siphon-sbc"   # per-role name so nodes appear separately in Homer

Putting it together

A solid baseline: scrape /metrics with Prometheus + alert on the table above; probe /admin/health + /admin/ready from your orchestrator; ship CDRs to your billing collector; and point HEP at Homer for call-flow forensics. For the production alert set and capacity guidance, see Deployment & operations.

See also