Monitoring & observability¶
You can see what SIPhon is doing four ways: Prometheus metrics (built-in + your own), the admin API, Call Detail Records, and full SIP tracing to Homer. None of them block the call path.
Prometheus metrics¶
Enable the endpoint:
SIPhon exports built-in gauges/counters; the ones worth alerting on:
| Signal | Alert when | Why |
|---|---|---|
siphon_memory_allocated_bytes |
rate(...[30m]) > 0 at flat call rate |
A real memory leak |
siphon_pyexec_jobs_shed_total |
sustained rate() > 0 |
Handler pool saturated → SIP retransmits |
siphon_pyexec_pool_size vs _pool_max |
pinned equal + all busy for minutes | Pool fully grown and saturated |
siphon_proxy_dialog_sessions |
grows under flat completed-call load | Dialog state not draining |
siphon_rtpengine_instances_up |
drops below your engine count | An RTPEngine is unhealthy |
siphon_diameter_peer_up{peer} |
== 0 |
That named peer is down. The bare siphon_diameter_peers_connected count cannot say which one |
siphon_diameter_answers_total{result_code!~"2.*"} |
rate() > 0 |
A peer is reachable and refusing. siphon_diameter_request_errors_total counts transport failures only, so this reads zero there |
siphon_ro_denials_total |
rate() > 0 |
Calls being refused credit by the OCS |
siphon_ro_credit_teardowns_total{reason="no_teardown_hook"} |
increase() > 0 |
Credit ran out with nothing wired to enforce it — the call is still up, and unpaid |
siphon_diameter_inbound_answers_total{result_code="3002"} |
rate() > 0 |
Server role only: siphon is rejecting inbound requests because no @diameter.on_request handler matched — a script gap, not a peer problem |
siphon_diameter_inbound_answers_total{result_code="5012"} |
rate() > 0 |
Server role only: an @diameter.on_request handler raised or returned the wrong type. Distinct from 3002 on purpose — on Ro a 3002 reads as a credit denial and tears the call down |
siphon_gateway_source_last_success_timestamp_seconds |
time() - <gauge> > 3 * refresh_secs, or == 0 |
The carriers are stale: siphon kept the last set it read and the controller has been unreadable since. 0 = never read, so a node that booted into a controller outage has no carriers from the source |
siphon_gateway_source_failures_total |
sustained rate() > 0 |
Reads are failing now; the timestamp above says for how long |
siphon_b2bua_inbound_calls_refused_total{reason} |
sustained rate() > 0 |
Inbound calls are being turned away at b2bua.inbound_limit. reason="concurrent" is the instance full, reason="rate" is a burst arriving faster than max_calls_per_second |
siphon_b2bua_calls_active / siphon_b2bua_max_concurrent_calls |
ratio above your headroom, e.g. > 0.8 |
Capacity is running out before calls are refused. The limit gauge reads 0 when no ceiling is set, so guard the division |
siphon_gateway_inbound_calls_refused_total{group,reason} |
sustained rate() > 0 |
A carrier is being turned away at its group's inbound_limit. Either it is sending more than was agreed, or the limit is set below what was |
siphon_gateway_inbound_calls_active{group} / siphon_gateway_inbound_max_concurrent_calls{group} |
ratio above your headroom | A carrier is close to its ceiling. Only groups with an inbound_limit have these series, and they are refreshed every 30 s |
siphon_registrant_source_last_success_timestamp_seconds / _failures_total |
same pair | Outbound trunk registrations, tracked separately — a node can route over stale carriers with current registrations, or the reverse |
See Handler execution model for the pool internals.
Your own metrics¶
The metrics namespace adds counters, gauges, and histograms that appear on the same
/metrics endpoint:
from siphon import metrics
calls = metrics.counter("calls_total", "Calls processed", labels=["direction", "result"])
active = metrics.gauge("calls_active", "Active calls", labels=["direction"])
setup = metrics.histogram("call_setup_seconds", "INVITE→200 latency",
buckets=[0.1, 0.25, 0.5, 1, 2.5, 5])
calls.labels(direction="outbound", result="ok").inc()
active.labels(direction="outbound").inc() # ... .dec() when it ends
setup.observe(0.342)
Admin API — health, readiness, registrations¶
A separate HTTP port for probes and runtime inspection:
| Endpoint | Use |
|---|---|
GET /admin/health |
liveness — 200 while the process is alive (survives drain) |
GET /admin/ready |
readiness — 200, or 503 while draining (SIGTERM) |
GET /admin/stats |
uptime + active registration count |
GET /admin/registrations[/{aor}] |
inspect bindings |
DELETE /admin/registrations/{aor} |
force-unregister |
GET /admin/bans / DELETE /admin/bans/{ip} |
list / lift auto-bans (404 on both when security.failed_auth_ban is off, so an empty list always means "watching, nothing banned") |
GET /admin/logs |
retained log lines (WARN+, plus down to admin.log_tail.retain_level); filters level, contains, call_id, paging limit + before=<seq> |
GET /admin/gateways |
per-group dispatcher status (destinations, health, weight, priority, missed health-checks) |
POST /admin/gateways/{group}/{destination}/{up\|down} |
mark a gateway destination up/down (drain / restore a carrier) |
POST /admin/registrants/refresh |
re-read a registrant.backend: database / http source and reconcile now |
POST /admin/gateways/refresh |
re-read a gateway.backend: database / http source and reconcile now |
GET /admin/calls |
active B2BUA calls (Call-ID, state, caller, callee, B-legs) |
GET /admin/metrics.json |
curated JSON snapshot of the live gauges + counters |
Point Kubernetes liveness at /admin/health and readiness at /admin/ready so a
draining pod leaves rotation cleanly — see Deployment & operations.
Bearer-token auth¶
The admin API can force-unregister bindings and lift bans, so gate it once it is reachable by anything but localhost:
admin:
listen: "127.0.0.1:9091"
auth:
token: "${ADMIN_TOKEN}" # keep the literal out of YAML
protect_reads: false # true = also require it on GET + /metrics
With a token set, the DELETE routes require Authorization: Bearer <token>
(constant-time compared). Reads stay open unless protect_reads is true. Unset
leaves the API open, exactly as before.
Web dashboard (experimental)¶
A single-page operator dashboard is baked into the binary and served same-origin on the admin listener. It's experimental — expect changes. The release Docker image compiles it in; you just enable it in config:
A plain cargo build leaves the ui feature off, so any project embedding
siphon as a library carries none of it; a binary built without --features ui
logs a warning and serves nothing when enabled is set. Serving the dashboard
logs an EXPERIMENTAL warning.
The views are Overview, Calls, Registrations, Gateways, Signalling, Media,
Control, Security and System, reading /admin/metrics.json,
/admin/registrations, /admin/calls, /admin/gateways and /admin/bans.
Force-unregister, lift-ban and gateway mark-up/down go through the same bearer
token — click Unlock and paste the admin.auth.token. Bind the listener
internally and put it behind your own ingress auth for anything beyond a trusted
network.
The dashboard is deliberately not a Grafana replacement. Rates and history belong in Prometheus; what the dashboard shows is the live, listable state Prometheus cannot hold — which calls are up, which AoR is bound where, which gateway is down, which control app owns which channel. You cannot put a Call-ID in a Prometheus label.
A subsystem this node has not configured shows as "not configured" rather than a zero, and its nav entry is greyed. That distinction matters: a permanent zero reads exactly like an idle subsystem, which is how a set of never-written metrics went unnoticed on this dashboard for a long time.
Reading the traffic counters¶
siphon_requests_total{method,direction} and
siphon_responses_total{class,direction} count wire events, not
transactions. Retransmit detection happens well downstream of the point they are
taken, so a UDP INVITE retransmitted by Timer A counts once per datagram, and a
non-2xx INVITE response retransmitted until ACK counts once per send. That is
the right meaning for a receive/send counter, but it means requests_total is
not a call count and not a transaction count — use siphon_dialogs_active (or
its siphon_proxy_dialog_sessions / siphon_b2bua_calls_active halves) for
concurrency, and CDRs for completed calls.
A method siphon does not recognise is counted under OTHER rather than by its
own name. The method is read straight off the request line, so labelling by it
would let any peer mint unbounded Prometheus series through the scrape endpoint.
siphon_connections_active{transport} has no UDP series. UDP is
connectionless, so there is no connection to count; a zero there would read as
"no UDP traffic", which is a different and wrong statement.
siphon_non_sip_datagrams_dropped_total{reason} is not an abuse counter and
feeds nothing into failed_auth_ban. It counts payloads siphon refused to hand
the SIP parser because they cannot be a SIP message: whitespace (RFC 3261
§7.5 / RFC 5626 §4.4.1, any transport), all_nul (a vendor NAT keepalive no RFC
defines, sent at a registered contact every few seconds for the life of the
binding), and too_short (a UDP datagram under the 14-byte grammar floor for a
SIP start line). A steady all_nul rate proportional to your registration count
is normal. What is worth an alert is a step change — a peer changed
behaviour — or a too_short rate with no keepalive behind it, which is worth one
RUST_LOG=siphon::dispatcher=trace session to see the bytes. Anything the parser
does reject still warns, but only once per source per minute, with the next line
carrying the count that went unlogged.
Call Detail Records¶
cdr:
enabled: true
auto_emit: true # write one CDR per call automatically
include_register: false # with auto_emit, also emit a CDR per REGISTER
backend: http # file | http | syslog
http:
url: "https://collector.example.com/v1/cdr"
auth_header: "Bearer tok123"
CDRs are written asynchronously (a bounded channel, never blocks a call) with the call's timing, parties, transport, disconnect initiator, and response code.
With auto_emit: true siphon writes one CDR per call on its own — proxy or B2BUA,
no script needed — filling in timestamp_start/answer/end, duration_secs,
response_code, and disconnect_initiator (caller / callee / timeout /
error). Answered calls, B-leg failures, answer timeouts and caller CANCELs all
produce a record. It defaults off, so it never surprises a manual-only setup.
You can also write records from a script — either instead of, or on top of,
auto_emit (use it to attach billing_id / trunk / account fields). cdr.write()
takes the proxy request or, from a B2BUA handler, the call:
from siphon import cdr
@proxy.on_request("INVITE")
def route(request):
cdr.write(request, extra={"billing_id": "B-12345", "account": "ACC-789"})
@b2bua.on_answer
def answered(call, reply):
cdr.write(call, extra={"billing_id": "B-12345"})
With auto_emit on, that does not write a second record: siphon is already
tracking one for the call, and extra is merged into it, so the fields land on
the same row as duration_secs, response_code and disconnect_initiator when
the call ends. Call it as often as you like — later fields merge on top, last
write wins per key. A record of its own is written only when there is nothing to
merge into: auto_emit off, or a request the auto-emit hooks do not track (a
MESSAGE, an out-of-dialog request).
So attach fields as soon as you know them (@proxy.on_request("INVITE"),
@b2bua.on_answer) rather than trying to assemble a record at teardown — the
teardown half is siphon's job.
A call the proxy never forwards (the script answered a 403, or dropped it)
normally produces no CDR at all. Calling cdr.write() on it is how you ask for
one anyway — a blocked-call record, emitted with the code the script answered.
Watch siphon_cdr_sessions (the live per-call tracking count): it returns to 0
between calls, and a steady climb under flat load means a call teardown isn't
being seen.
Full SIP tracing → Homer¶
Stream every SIP message to a Homer / heplify-server collector over HEP — invaluable for debugging call flows:
tracing:
hep:
endpoint: "127.0.0.1:9060"
version: 3
transport: udp # udp | tcp | tls
agent_id: "siphon-sbc" # per-role name so nodes appear separately in Homer
Putting it together¶
A solid baseline: scrape /metrics with Prometheus + alert on the table above; probe
/admin/health + /admin/ready from your orchestrator; ship CDRs to your billing
collector; and point HEP at Homer for call-flow forensics. For the production alert
set and capacity guidance, see Deployment & operations.
See also¶
- Real example:
examples/timer_example.py(periodic health pushes),siphon.yaml. - Deployment & operations — the ops runbook.