SIPhon Feature Readiness Matrix¶
Overview¶
This document tracks the maturity of every SIPhon feature across three readiness levels. SIPhon runs in production today in a residential SIP registrar/proxy role and a 3GPP IMS deployment exercising Diameter Cx/Sh/Rx, iFC, IPsec, and 5G SBI policy control. Features validated on live traffic are marked Production.
| Readiness | Meaning | Evidence behind it |
|---|---|---|
| Production | Running on live traffic today | Everything below, plus a deployment carrying real calls |
| Implemented | Code-complete and tested, not yet production-deployed | At least unit/integration tests; often a SIPp scenario too — but see below, the level does not say which |
| Planned | Partially wired or design-only | None yet |
| Not implemented | Named here because its absence is worth knowing | n/a |
What these levels do not tell you¶
"Implemented" spans a wide range of evidence. A row backed by unit tests
only and a row backed by a SIPp scenario that runs on every pull request both
read Implemented. When a row's evidence matters to you, the Notes column
names it — most rows say which tests exist — but the level alone will not.
The gates, weakest to strongest:
| Gate | What it proves | When it runs |
|---|---|---|
| Unit / integration tests | The code does what its author intended | Every PR |
| RFC 4475 corpus + parser proptests | Adversarial messages parse or are refused per the RFC | Every PR |
| Fuzz targets | The parser and stream framer survive hostile input, and agree with each other | Every PR (60 s each) |
| SIPp scenarios | The feature works end to end against a message generator | Most on every PR; some only via scripts/run-tests.sh |
| Throughput + memory-leak baseline | 16 rows of proxy/B2BUA × UDP/TCP hold with zero failures and flat allocation | Release cut |
| Live traffic | It survives real endpoints | Production rows only |
No level asserts interoperability with an independent implementation. SIPp
is a message generator, not a SIP element: it builds no route set, matches no
CANCEL to a transaction, and has no opinion about a Record-Route it did not
write. A feature can pass every gate above and still be wrong in a way only
another stack would notice. Closing that is what interop/ is for; until it
covers a feature, read Implemented as "correct by siphon's own account".
Every SIPp scenario now has a runner. Four previously had none —
b2bua-refer-outbound, b2bua-reinvite-bleg, b2bua-reinvite-breject,
b2bua-reinvite-reject — so the behaviour they described was asserted by
nothing. Running them for the first time found one real defect (a
siphon-originated REFER emitted ahead of the A-leg 2xx) and one scenario bug.
All four, plus the reliable-provisional scenario that had been driven only by
scripts/run-tests.sh, now run on every pull request.
Core SIP Engine¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Stateful proxy (RFC 3261 §16) | Production | script: @proxy.on_request |
Full transaction state machines; ICT Timer A RFC-compliant (capped at T2, fires in Proceeding, cancelled on final response) |
| B2BUA (RFC 3261 §6) | Production | script: @b2bua.on_invite |
Two-leg call control, per-leg Call-ID + From-tag, topology hiding. The call.fork/call.dial timeout= (default 30s) is now enforced: a B-leg INVITE sent fire-and-forget (no client transaction, so no Timer B) that never produces a final 2xx — dead/partitioned trunk — is failed within timeout..timeout+0.5s by the dedicated 500ms answer-timeout check (check_b2bua_answer_timeouts → fail_b2bua_call_on_timeout; moved off the 30s orphan sweep so a short per-carrier LCR ring timeout re-routes promptly): CANCEL pending legs, @b2bua.on_failure(408), then whatever it decides (by default 408 Request Timeout to the A-leg and teardown). Previously the call leaked until the 24h orphan backstop. Unit-tested in take_timed_out_calls_only_unanswered_past_deadline. Outbound auth-retry 2xx ACK: a B-leg INVITE to an authenticating trunk (401/407 → credentialed CSeq-2 retry) is superseded in place (replace_b_leg), which drops the failed leg's actor handle so that actor emits CallEvent::Terminated onto the SHARED per-call classification channel. The dispatcher block-recvs that channel per response; consuming the stale Terminated desynced the stream so the retry leg's 200 OK was misclassified as the prior 18x's provisional — set_winner and the B-leg ACK were skipped, so the trunk's 200 OK retransmitted unacked and the call collapsed into a BYE storm ~5 s after answer (all outbound PSTN to an authenticating trunk). Fixed by recv_b_leg_classification_event skipping Terminated lifecycle events when reading a response classification (regression-tested in dispatcher::tests::b_leg_200_classifies_as_answered_after_auth_retry_supersede). Outbound retry member affinity (RFC 5923): the 401/407 credentialed re-INVITE and the RFC 4028 422 higher-Session-Expires re-INVITE are fresh pre-dialog transactions (new branch + CSeq, no To-tag), so the in-dialog connection-reuse path did not cover them — they re-resolved the trunk hostname and the RFC 3263 §4.2 A/AAAA shuffle could land the retry on a different member of a multi-member trunk (one DNS name) than the one that issued the nonce, drawing a second 401 on a strict trunk (auth loop) or splitting one INVITE transaction across two members (fragile CANCEL/BYE/session-timer correlation). Both retries now reuse the failed leg's established destination/transport/connection_id (select_b2bua_retry_destination) so the whole transaction stays on the nonce-issuing member; falls back to fresh DNS resolution only when the leg has no recorded destination (regression-tested in dispatcher::tests::b2bua_retry_reuses_established_member_not_resolved_sibling). A B-leg response relayed to the caller (an 18x, the 2xx, a relayed failure) carries the caller's own From and To as they arrived, with the A-leg dialog's tag (RFC 3261 §8.2.6.2), never siphon's B-leg identity as a number policy or rewrite_identities() shaped it, nor the far side's address; failures siphon builds itself already did. Previously a relayed failure swapped only the tags. Unit-tested in dispatcher::relayed_identity_tests; SIPp-validated by the sipp-b2bua-lcr CI job (an LCR sequence whose last carrier answers 183, then 404). A 101-199 on a B-leg that has already ended (an LCR carrier settled by its failure, a CANCELled leg, a fork branch that failed) is dropped: it is not relayed to the caller, records no progress and runs no @b2bua.on_early_media. Dispatcher-driven tests in dispatcher::late_provisional_tests. The callee's 2xx is ACKed as soon as it has been handled, and so is every retransmission of it, whether or not the caller has ACKed yet (RFC 3261 §13.2.2.4). The ACK used to be held until the caller's ACK arrived, with the retransmissions absorbed, so a caller whose ACK was late or lost left the callee unACKed for all of 64T1. Every re-ACK now carries the dialog's route set as well. Dispatcher-driven tests in dispatcher::b_leg_2xx_ack_tests; SIPp-validated by the sipp-b2bua CI job (a caller that holds its ACK for 4 s while the callee needs siphon's within 2 s). Delayed offer (RFC 3264 §4): when the B-leg INVITE went out without an offer, the callee's 2xx carries it and the ACK is held (CallActor::delayed_offer_ack) until the caller's ACK brings the answer, which it then carries with what any relayed SDP gets (siphon's o=/s=, the leg's o= version, media.sdp_strip_attributes); copies of the 2xx are absorbed while it waits and re-ACKed with the same ACK after. A call ending before the caller answers ACKs the callee with every stream rejected (port 0) right before its BYE, and a caller ACK with no answer ends the call that way (Reason: Q.850;cause=111). Media-anchored too: rtpengine.answer(reply) sends a 2xx that carries the offer to the engine as an offer from the callee's side, and send_delayed_offer_ack sends the caller's answer as the engine answer and ACKs the callee with the engine's SDP; an engine refusal ends the call with Q.850;cause=47 (tests in delayed_offer_ack_tests and script::api::rtpengine::answer against an in-process NG engine, rtpengine::test_engine). Dispatcher-driven tests in dispatcher::delayed_offer_ack_tests; SIPp-validated by the sipp-b2bua CI job (an offerless caller whose callee must receive an ACK carrying the answer). A caller that never ACKs the 2xx siphon sends it, relayed or siphon's own (call.answer(), the control plane's answer), has the call ended after 64T1 (RFC 3261 §13.3.1.4): sweep_unacked_uas_2xx on the 100 ms timer tick BYEs both legs through the framework teardown with Reason: Q.850;cause=102;text="No ACK received". The retransmit task used to log and leave the call up. An ACK processed first always wins, since the sweep only acts on a retransmit entry it removes itself, and a call already torn down gets no second BYE. Paused-clock dispatcher tests in dispatcher::unacked_answer_tests; SIPp-validated by the sipp-b2bua CI job (a caller that never ACKs, both legs must receive the BYE; about 33 s, since T1 is not configurable). A call that ends before the caller has ACKed its 2xx (the callee hangs up, or any teardown through b2bua_terminate_call_inner) BYEs the callee at once and holds the caller's BYE in held_a_leg_byes (RFC 3261 §15): the 2xx keeps being retransmitted, the caller's ACK sends the BYE right after it, and at 64*T1 sweep_unacked_uas_2xx sends it once. The caller's 2xx is registered as waiting for its ACK before the callee's ACK or the 2xx itself goes out, so a callee that BYEs the moment it is ACKed, handled on another worker while siphon is still sending the caller its 2xx, is held the same way (forced interleaving in dispatcher::b_leg_2xx_ack_tests). A caller BYE in that window is answered 200 and drops the held BYE. Every BYE siphon sends goes through send_or_hold_bye, which decides by the dialog's Call-ID (the 2xx store is keyed by it), never by leg slot, so the transfer BYEs (REFER referrer, immediate or after its NOTIFY; survivor of a failed transfer; the party a Replaces takes over) are held the same way, and a caller a takeover moved into the B-leg slot is too; dispatcher tests in dispatcher::transfer_bye_tests. Teardowns are claimed per call (CallActorStore::claim_teardown): of two that start together, only the first sends BYEs, the others back off and a BYE arriving meanwhile is answered 200; forced interleavings in dispatcher::teardown_race_tests. Dispatcher-driven tests, most on a paused clock, in dispatcher::held_bye_tests; SIPp-validated by the sipp-b2bua CI job (a callee that hangs up right after its ACK and needs its BYE answered within 1.5 s, while the caller holds its ACK for 3 s and fails on any BYE before it). On a media-anchored call a re-INVITE or UPDATE relayed between the two legs pins each party's media ingress (received_from) by its own half of the profile, the caller's offer half and the callee's answer half, whichever of them re-offers, on the re-offer and on the answer; the commands keep their shape (dispatcher::reoffer_ingress_tests a_relayed_reoffer_pins_each_party_by_its_own_half_whoever_reoffers, read off the commands an in-process native engine records). On an anchored delayed offer the caller's answer in the ACK is pinned by the caller's offer half to the caller's signalling source, and the answered media session names the caller first, so a re-INVITE from either party afterwards is re-offered under that party's own tag (dispatcher::delayed_offer_ingress_tests). |
| Parallel forking | Production | request.fork() (proxy); call.fork() (B2BUA) |
Proxy: used for AS→subscriber delivery. The branches that lost to a 2xx or a 6xx are CANCELled once each has drawn a provisional and none that already has its final response (RFC 3261 §9.1, §16.7 step 10), and a provisional arriving after the fork settled is not forwarded (§16.7 step 5): proxy::fork::tests::a_provisional_after_the_fork_settled_is_not_forwarded, dispatcher::proxy_cancel_awaits_provisional_tests. B2BUA: the first branch to answer wins and the rest are CANCELled on it (RFC 3261 §9.1); their 487s are ACKed without touching the answered call, and a 2xx that crosses the CANCEL is ACKed and BYEd (§13.2.2.4, §15). A branch failing leaves the caller ringing while another can still answer; once none can, the caller gets the best of the branches' failures and @b2bua.on_failure runs once (§16.7). Proxy and B2BUA choose that failure with one shared ranking (§16.7 step 6: any 6xx, else the lowest class, 401/407/415/420/484 preferred within 4xx, 503 last within 5xx and sent to the caller as a generated 500); the proxy forwards the chosen branch's own response with every other 401/407 challenge added (step 7). Unit-tested in sip::best_response and proxy::fork. A 6xx settles at once and CANCELs the rest. On the ring timeout, a failure a branch already returned is relayed instead of a bare 408 whenever it outranks one (§16.8), and the branches still ringing are CANCELled. Branches are sent one at a time, so settlement waits until all are out: a branch failing (or answering) before its siblings exist neither fails the call nor leaves them ringing. Previously the first failed branch failed the whole call and losing branches were never CANCELled. Unit-tested in b2bua::actor::tests (settlement, dispatch window, cancel-on-win); SIPp-validated by the sipp-b2bua-fork CI job (busy + answer, answer + cancelled loser, answer + loser 2xx glare, every branch failing, busy + ring timeout, busy + 503, a lone 503 sent up as 500, busy + redirect keeping the callee's Contact, a sequential ring-out + 503 sending the 408). |
| RFC 2543 legacy transaction matching | Not implemented | — | RFC 3261 §17.2.3 also defines matching for a request whose topmost Via carries no branch, or one without the z9hG4bK magic cookie (Request-URI + To tag + From tag + Call-ID + CSeq + top Via, with ACK keyed on the To tag of the response). siphon does not implement it: key_from_message fails and no server transaction is created. The request is still processed, statelessly — so retransmissions are not absorbed and each one runs the script again. Counted as siphon_requests_without_branch_total; a non-zero rate means a legacy or broken peer is on the network. A branch has been mandatory since RFC 3261 §8.1.1.7 (2002) and nothing in siphon's deployment profile — IMS, WebRTC, modern trunks — speaks RFC 2543, so this is a documented gap rather than a half-built implementation (the ACK case needs a second index keyed on a tag siphon only learns when it sends the response) |
| Sequential forking | Implemented | request.fork(strategy="sequential") (proxy); call.fork(strategy="sequential") (B2BUA) |
Proxy: fork-aggregator TryNext. The proxy moves on only when the branch it is on has its final response or its INVITE has timed out (Timer B, answered as a 408), so the branch it leaves is never CANCELled (dispatcher::proxy_cancel_awaits_provisional_tests::a_sequential_fork_moves_on_without_cancelling_the_branch_it_leaves, a_failure_retarget_does_not_cancel_the_branch_that_failed). B2BUA: previously the strategy was silently ignored (every target rung in parallel); now honored via the sequential-failover engine (shared with LCR call.route) — targets tried one at a time, each a fresh B-leg dialog, advancing on failure. |
| Least-Cost Routing (LCR) | Implemented | lcr: + await lcr.route(call) + call.route(decision.routes) (B2BUA-only) |
The external HTTP JSON API owns the cost/order decision (siphon is not a rating engine); siphon caches it (named cache, fleet-wide via Redis) with a static fallback group, and executes the ordered route set with sequential failover: cheapest-first, each carrier a fresh B-leg dialog (new Call-ID — no reused Call-ID across carriers, the proxy serial-fork footgun), resolving a route's gateway_group to a healthy member (skip a dead pool) and advancing on a reroute cause. A route's timeout_secs bounds the wait for progress (a 101-199, the RFC 3261 §16.7 step 2 Timer C line): a carrier that has shown none by then is CANCELled and failed over, while one that has keeps the call to the later of its own timeout and the sequence's ring bound (call.route(timeout=…)), then the call fails 408 without advancing; per-route reroute_after_progress: true restores failover for a carrier that fakes progress with its own ringback. Sequential forks and a sequential control-plane dial still advance per target (hunt semantics). Dispatcher-driven tests in dispatcher::lcr_ring_timeout_tests (UDP egress read back, answer sweep stepped past each deadline), and SIPp-validated in the sipp-b2bua-lcr CI job (two carriers with timeout_secs: 3 under an 8 s ring bound: the first answers 183 and is CANCELled at the ring bound, not at its 3 s, the second never receives an INVITE, and the caller gets one 408 between 7 and 10 s after its INVITE, ACK absorbed). A carrier's leg is settled by its final failure, so no later ring timeout, caller CANCEL, answer or @b2bua.on_failure re-route CANCELs a carrier that already answered (RFC 3261 §9.1), and a retransmitted failure is re-ACKed without touching the sequence. Every ring timeout, the last carrier's and a progress-kept carrier's included, is recorded once as 408 on call.route_attempts / lcr_attempts and fires @b2bua.on_route_failure once, before @b2bua.on_failure. A sequence that ends on the ring timeout of a carrier that never sent a 101-199 fails the call with siphon's own 503 (not rewritten to 500), and one whose carrier had rung still fails 408, including the last target of a hunt; a sequential control-plane dial reports the same code in DialFailed. Dispatcher-driven tests in dispatcher::lcr_route_bookkeeping_tests, and SIPp-validated in the same sipp-b2bua-lcr job (a carrier that answers 503 and then sends a stale 183 on that dialog, and a last carrier with timeout_secs: 3 that answers only 100 and is CANCELled at its 3 s: the stale 183 never reaches the caller, the first carrier receives no CANCEL in the 10 s after it, the caller gets one 503, and siphon logs the 503 and then the ring timeout's 408 once each through @b2bua.on_route_failure, before its one @b2bua.on_failure with 503). Reroute causes selectable per-route (API) > per-gateway (gateway.groups[].reroute_causes) > global (lcr.reroute_causes, default [408,500,502,503,504]); a definitive response (486/603) is forwarded to the caller, and an exhausted sequence sends the best of its carriers' failures (RFC 3261 §16.7), not the last. Per-carrier number_policy (shapes the dialled number in the R-URI and the identity headers alike; a route naming none gets b2bua.default_number_policy), tech_prefix (prepended to the shaped R-URI number), full ruri override, injected headers. Per-carrier presented CLI and CLIR: caller_id substitutes the calling number on From and on PAI/PPI through the tag-preserving identity path, and caller_id_presentation: "restricted" withholds it per RFC 3323 §4.1 / 3GPP TS 24.607 — anonymous From with the tag intact, Privacy: id appended, PAI kept for the trusted next hop, PPI removed — the four moving together so a carrier that renders From rather than PAI cannot leak the number. A copied Remote-Party-ID moves with them: rewritten by caller_id, removed by CLIR and by a controller dial identity that replaces the caller's (sip::privacy unit tests, dispatcher::control_dial_rpid_tests). A PAI is asserted when the leg has none (b2bua.assert_identity, on by default), which on a B-leg is the usual case: the header policy strips P-* at the trust boundary, so caller_id used to rewrite a From and nothing else, and Privacy: id used to be asserted with nothing behind it. Ordering is caller_id → assert → number policy → CLIR, each step depending on the one before it: the PAI carries the presented number rather than the caller's own, agrees with the From's format, and on a restricted route naming no caller_id carries the real identity, asserted at the last point the From still holds it. Unit-tested in sip::privacy::tests and dispatcher::b_leg_asserted_identity_tests (read off the UDP egress, since the ordering is the whole of it), and SIPp-validated in the sipp-b2bua-lcr job by a carrier asserting on the PAI, the anonymised From and the Privacy header it was actually sent. Both are per-route with no answer-level default, so a failover never inherits the previous carrier's presentation; an unrecognised value is logged and treated as restricted. Script-level twins call.set_caller_id() / call.restrict_caller_id() for non-LCR deployments. call.active_route surfaces the winner for CDR. B2BUA-only by design (dialog hygiene / per-carrier media / charging). Rust unit tests (contract + client mock-HTTP cache/fallback/reject + actor failover state), SDK-mirrored + tested (sdk/tests/test_lcr.py), end-to-end SIPp-validated locally (scripts/lcr_sipp_test.sh: carrier-A 503 → transparent failover → carrier-B 200, caller sees one 200) and in CI by the sipp-b2bua-lcr job (docker, mock LCR API, policy-shaped carrier INVITEs: carrier-A 503 fails over, carrier-B's 183 and then 404 reach the caller with its own From and To, and its ACK is absorbed). Charging (Rf/Ro carrier stamping) lands with online charging. |
| B2BUA inbound limit | Implemented | b2bua.inbound_limit |
A ceiling on the inbound calls an instance accepts: max_concurrent_calls and max_calls_per_second (GCRA, burst of one second's worth), both unset by default. Checked on the inbound initial INVITE before @b2bua.on_invite; a call past either is answered 503 with Retry-After (RFC 3261 §21.5.4; reject_code / retry_after_secs), statelessly, and no handler fires. Calls siphon originates, emergency calls (urn:service:sos, RFC 5031) and Replaces takeovers hold a slot and are never refused. The slot is a permit owned by the call actor, so every teardown path releases it by dropping the call. A retransmitted refused INVITE is re-answered from a bounded table and counted once (RFC 3261 §17.2.1). Each refusal writes a CDR (disconnect_initiator="local", refusal_scope, refusal_reason) and counts in siphon_b2bua_inbound_calls_refused_total{reason}; the limits are published as gauges. Per instance. Unit-tested in admission (exact ceiling under contention, the GCRA schedule against hand-computed times, drain to baseline) and through the dispatcher in dispatcher::admission_tests (refusal is the only frame on the wire; the slot returns after a callee failure, a caller CANCEL, a ring timeout, an answered call's BYE, a script reject and a silent drop). SDK-mirrored (SipTestHarness.set_inbound_limit, CallDetailRecord.is_refused). SIPp-validated by the sipp-b2bua-admission CI job, which also asserts from the script's log that the refused call never reached it. Criterion bench admission. |
| Gateway group inbound limit | Implemented | gateway.groups[].inbound_limit / inbound_* row fields |
The b2bua.inbound_limit block on a gateway group, capping calls arriving from the sources the group admits (resolved destination addresses plus source_networks, the from_gateway() membership). B2BUA only. Judged before the instance limit, so a carrier over its own is refused with its group's reject_code and costs the instance nothing; a source two limited groups admit is counted against both. The counters are kept by group name outside the group object, so a provisioning reconcile that rebuilds the group keeps the count and a changed limit applies to the calls already up; a removed group's count is forgotten. The per-INVITE lookup reads a lock-free snapshot of only the limited groups (one atomic load when there are none; about 19 ns per limited group on the reference machine). Provisioned groups carry it through inbound_max_concurrent_calls / inbound_max_calls_per_second / inbound_reject_code / inbound_retry_after_secs; script-added groups cannot. CDR refusal_scope: gateway + gateway_group; siphon_gateway_inbound_calls_refused_total{group,reason}, siphon_gateway_inbound_calls_active{group} and the limit gauges (30 s sweep); inbound_limit in GET /admin/gateways. Unit-tested in gateway::inbound_limit_tests (overlap, rebuild keeps the count, remove and re-add, churn drains to baseline), gateway::source::tests, and through the dispatcher in dispatcher::admission_tests. SDK-mirrored (GatewayRow, MockGateway.set_inbound_limit, CallDetailRecord.refusal_gateway_group). SIPp-validated by the sipp-b2bua-admission-gateway CI job. |
| Record-Route / Loose Route | Production | request.record_route() |
Mid-dialog routing proven. A relay stamps one entry per socket the dialog crosses, not per transport, so a proxy bridging two same-transport listeners (a P-CSCF joining its protected access port to its core port) tells each peer the socket facing it. A UAS answering a dialog-forming request echoes every entry back per RFC 3261 §12.1.1 — automatic in the response builder, so it covers request.reply() and call.answer() alike; scoped to an initial INVITE/SUBSCRIBE/REFER answered 101–299, and copied verbatim so an unknown parameter survives. Overridable with set_reply_header("Record-Route", …) |
| CANCEL propagation | Production | Core | Matched to transaction automatically. Every proxy CANCEL is built by the branch's INVITE client transaction from the INVITE as it sent it (RFC 3261 §9.1/§16.10: the same Request-URI, Route set, From, To, Call-ID and CSeq number, and that INVITE's top Via, branch and sent-by, as its one Via), so the downstream proxy/UAS matches CANCEL→INVITE (§17.2.3) and tears the alerting branch down; a Reason in the caller's CANCEL is relayed (RFC 3326). This replaced a CANCEL assembled from the caller's request, which had the caller's Request-URI and route set rather than the branch's, and before that one with a fresh branch that matched nothing (transaction::cancel_tests::the_cancel_is_built_from_the_invite_as_it_was_sent, and every CANCEL in dispatcher::proxy_cancel_awaits_provisional_tests is compared with its INVITE). B2BUA CANCEL path (build_cancel_from_invite) builds the correct per-leg branch; additionally fixed a defect where a 401/407 digest or RFC 4028 422 retry on an outbound INVITE appended a fresh B-leg instead of superseding the failed one, so a caller CANCEL during alerting fanned out to the dead pre-auth transaction too (→ a spurious 481, RFC 3261 §9.1). Retries now replace the leg in place (CallActorStore::replace_b_leg), so CANCEL targets only the live branch (regression-tested in b2bua_auth_retry_supersedes_failed_leg_for_single_cancel). 2xx-after-CANCEL glare (RFC 3261 §9.1): when the callee answers a B-leg INVITE in the cancel window, the B2BUA used to drop the racing 200 OK as an unknown branch (the call was already removed), leaving the callee retransmitting 200 OK then BYEing the half-open dialog. handle_b2bua_cancel now preserves still-pending B-legs as zombie_cancelled entries (32 s window); the racing 2xx is ACKed (§13.2.2.4) and immediately BYEd (§15) by handle_zombie_cancelled_2xx. A B2BUA CANCEL also waits for the INVITE's first provisional response (§9.1: "the CANCEL request MUST NOT be sent" before one): giving up on an INVITE that has drawn no response leaves it retransmitting, and its first provisional (a 100 Trying counts) draws the CANCEL, a 2xx an ACK and a BYE, any other final an ACK alone, and Timer B nothing at all. Timer B runs on a reliable transport too (§17.1.1.2), where nothing retransmits the INVITE (B2buaRetransmits::arm_timeout; cancel_awaits_provisional_tests::a_silent_branch_on_a_reliable_transport_ends_at_timer_b_without_a_cancel, b2bua::retransmit a_reliable_invite_times_out_at_timer_b_and_never_retransmits), and it ends only the wait for the CANCEL: the branch stays answerable until its expiry, so a 2xx after Timer B is ACKed and BYEd (a_cancelled_dial_to_a_phone_that_never_responds_ends_at_timer_b_without_a_cancel, timer_b_ends_the_wait_for_a_cancel_and_leaves_the_branch_answerable). One sender (cancel_kept_branches) serves every path that gives up on an INVITE siphon sent: originate and originate groups, bridging and connecting dial, fork branches, LCR failover, the caller's relayed CANCEL, drop / terminate / shutdown, and transfer or replace_peer targets. Unit-tested in b2bua::actor::store::cancelled (the wait, the single claim, each of the four ways it ends, and the kept branch draining to zero) and, read off the UDP egress, in dispatcher::cancel_awaits_provisional_tests (a fork's losing branch), dispatcher::originated_cancel_awaits_provisional_tests (a bridging dial given up with cancel_dial) and dispatcher::relayed_cancel_awaits_provisional_tests (a caller's CANCEL), each for a late provisional, a late 2xx, a late failure and no response up to Timer B, plus a callee that already failed not being CANCELled, replacement_fork_tests::a_losing_target_that_has_not_responded_is_cancelled_on_its_first_provisional and lcr_ring_timeout_tests::a_carrier_that_never_responded_is_failed_over_at_its_ring_timeout; SIPp-validated by the cancel-late-ringing and cancel-late-answer cases of the sipp-control-transfer job. The proxy path follows the same rule through each branch's INVITE client transaction, whose state is what the branch has drawn (cancel_proxy_branch, the one sender for a fork settled by a 2xx or a 6xx, reply.reject() and the caller's CANCEL): a branch in Proceeding is CANCELled at once, one in Calling keeps its CANCEL with the transaction until its first provisional while its INVITE goes on retransmitting, and one with its final response gets none; a final response or Timer B drops a waiting CANCEL unsent, and it survives the proxy session, which the caller's CANCEL removes at once. A request other than INVITE is not CANCELled. Unit-tested in transaction::cancel_tests (the wait, the single claim, each of the four ways it ends, a reliable transport ending at Timer B, the waiting count draining to zero, and a response racing the request on two threads) and, read off the UDP egress, in dispatcher::proxy_cancel_awaits_provisional_tests (each reason to CANCEL against a branch that rang, one that is silent and one that failed; the silent one followed to a late 100/180, a late 2xx, a late failure and Timer B; a TCP branch; the two races through the dispatcher; and the drain). A 2xx from a branch the proxy gave up on, whether it crossed the CANCEL or arrived with no provisional before it, is forwarded to the caller (RFC 3261 §16.7 steps 5 and 9), who ACKs it and ends its dialog (§13.2.2.4): no reply handler runs for it, it is not the call's answer in a CDR or on Rf, and its BYE does not close the call's record. Tested for a settled fork (2xx and 6xx), reply.reject() and the caller's CANCEL in dispatcher::proxy_late_answer_tests (the 2xx at the caller with the caller's Via stack, the caller's ACK and BYE at the callee, the BYE's 200 back, the record untouched, the stores drained), with the aggregator's rule in proxy::fork::tests::test_parallel_late_2xx_from_cancelled_branch_is_forwarded and a_retransmitted_2xx_is_not_another_answer. A 2xx to an INVITE that matches nothing the proxy still holds (the retransmission of an answer, an answer on a branch already timed out) is forwarded statelessly by its Via stack (§16.7 step 9) when its top Via is this instance's own and a second Via names the sender, with the framework's rewrites only and no reply handler (dispatcher::proxy_stateless_2xx_tests: a lost 200 repeated over UDP, an answer after the session is gone, a caller on TCP, and foreign or single-Via strays dropped). Limitation: siphon stamps no received/rport on a forwarded request's Via, so such a copy goes to the address the caller wrote there. Limitation: a proxied call's CDR and Rf session are kept per call (Call-ID and the caller's tag), not per dialog, so a dialog opened by a late 2xx is never recorded on its own, and a caller that keeps it and ends the first dialog closes the call's record with the first dialog's BYE. A call ended by the caller's CANCEL or by reply.reject() has its record written then (dispatcher::proxy_cancel_upstream_tests::a_cancelled_or_rejected_call_closes_its_record). |
| No-handler fallback (405 / 481 / OPTIONS) | Production | server.auto_options |
A request no @proxy.on_request handler claims: OPTIONS is answered 200 with Allow + Contact (RFC 3261 §11.2, or silence with auto_options: false); a method siphon does not implement is 405 + Allow (§8.2.1); an in-dialog request (To-tag) for an implemented method is 481 (§12.2.2), since what is missing is the dialog, not the method. That last case is what a late BYE reaches once a torn-down B2BUA call ages out of the 32 s torn-down set, e.g. a peer's BYE 32 s after its own CANCEL, which used to be 405. Dispatcher-driven tests in dispatcher::stale_in_dialog_request_tests, unit tests on build_no_handler_response, and the nohandler probe (scripts/options_fallback_test.sh) against the real binary in both auto_options modes. |
| In-dialog sequential routing | Production | request.loose_route() |
End-to-end 2xx ACK follows the dialog route set (top remaining Route after self-consumption, else R-URI), not the cached INVITE next-hop — correct through non-Record-Routing hops (transparent iFC AS, I-CSCF). RFC 5923 connection reuse: when the route-set next hop still resolves to the peer the dialog was established with, in-dialog requests (B2BUA BYE/re-INVITE/UPDATE/PRACK/2xx-ACK, proxy 2xx-ACK) keep the established connection/address instead of re-resolving — so a load-balanced trunk behind one DNS name (load-balanced Record-Route) is not re-shuffled (RFC 3263 §4.2) onto a sibling member that holds no dialog state; still resolves fresh for a genuinely divergent next hop. Validated proxy/B2BUA × UDP/TCP, 0 failures/retransmits |
| UAC-originated pre-loaded Route | Production | proxy.send_request(headers={"Route": "<sip:host;lr>"}) |
Next-hop selection for a script-originated out-of-dialog request now follows RFC 3261 §8.1.2 / §16.4: when the headers carry a Route (a pre-loaded route set) and no explicit next_hop, the request is sent to the first Route entry's ;lr loose-route target — the R-URI stays in the Request-Line and the Route rides along. Previously the Route was carried but ignored, and the destination was always resolved from the R-URI's home domain, so a script pre-loading the serving S-CSCF (e.g. MMTel-AS reg-event SUBSCRIBE/refresh/UN-SUBSCRIBE) took an extra I-CSCF hop + Cx LIR/LIA per operation. Precedence: explicit next_hop > first Route URI > R-URI. Regression-tested end-to-end (send_request_python_kwargs_preserve_body_and_content_type scenarios 5–6) + unit-tested (resolve_send_target_*, route_next_hop_*, parse_first_route_uri_*) |
| Call transfer (REFER, RFC 3515) | Implemented | B2BUA @b2bua.on_refer |
Three modes (terminate / transparent / siphon-originated); no handler registered → 603 Decline locally, so an in-dialog REFER on a tracked call is never blind-relayed back into the B2BUA. Capability advertisement is load-bearing and now complete: Allow carries REFER/NOTIFY and Supported carries replaces (RFC 3891 §6.2) on the A-leg 2xx, the B-leg INVITE and the REFER 202 — a transferor reads both to choose its transfer method (RFC 5589 §7.3), and withholding either silently downgrades a consultative transfer to a blind REFER whose Refer-To carries no Replaces, stranding the transferor's consultation dialog. The subscription NOTIFYs carry Event: refer;id=<REFER CSeq> (RFC 3515 §2.4.6) so two transfers on one dialog stay distinguishable, and the 202 carries the Contact RFC 3515 §2.2 makes mandatory on a subscription-creating 2xx. Inbound INVITE with Replaces (the takeover half) is gated on b2bua.accept_replaces, off by default. End-to-end SIPp scenarios (b2bua-refer, -reject, -terminate, -terminate-busy, -referrer-bye, -attended, -outbound) run on every PR and assert the advertisements and the id on the wire. On a siphon-terminated transfer the target's 2xx is ACKed only once the target is promoted and the subscription cleared: ACKed earlier, a target that hangs up as soon as the ACK lands had its BYE answered 481 (no dialog leg yet) and the surviving party was never released — 3 runs in 10 under load. The survivor's 200 to the re-INVITE that re-points its media can arrive after that target has already hung up and the call is gone; it is ACKed from the response itself (RFC 3261 §13.2.2.4, RFC 5407 §3.1.3), where it used to be dropped and left the survivor retransmitting it (unit-tested in dispatcher::b2bua::response::late_ack::tests). |
| Cancel teardown hook | Implemented | @proxy.on_cancel / @b2bua.on_cancel |
Fires once when a relayed (proxy) or B2BUA INVITE is CANCELled before a final response (RFC 3261 §9) — the only script teardown signal for a cancelled-before-answer call, which neither on_reply/on_failure (proxy: the 487 is generated at the transaction layer, never reaching a reply handler) nor on_bye (b2bua: no dialog was ever established) deliver. Receives the original INVITE (proxy fn(request)) / the Call (b2bua fn(call)); fire-and-forget, does not gate the 487. Exists to release per-call resources no BYE will ever clear (Diameter Rx/N5 QoS sessions, rtpengine media anchors). The B2BUA hook fires only in Calling/Ringing, so a 2xx that wins the cancel/answer glare (independently ACK+BYE'd by handle_zombie_cancelled_2xx) never triggers it — no answered call is torn down. Engine-registration unit tests (script::engine::tests::{proxy,b2bua}_on_cancel_decorator_registers_handler) + SDK dispatch tests (sdk/tests/test_on_cancel.py). |
| Reply-time proxy reject | Implemented | reply.reject(code, reason) (in @proxy.on_reply) |
Fail an in-progress proxied INVITE from the reply context — the proxy-side equivalent of B2BUA call.reject(), needed because IMS P-CSCF media authorization (N5 sbi.create_session / Rx diameter.rx_aar) runs at answer time, when the negotiated SDP is available, and a failure must reject the leg (e.g. 503) rather than proceed medialess. On a provisional (1xx) — typically a reliable 183 in the VoLTE preconditions / early-media flow where the SDP answer rides the provisional — records the reject and returns True; the dispatcher then sends code reason upstream to the UAC via the server transaction (retransmission + UAC-ACK absorption) and CANCELs every pending downstream branch (reusing cancel_fork_branches, RFC 3261 §9). The straggler 487 the CANCEL draws back is absorbed via a new ProxySession.final_response_sent guard (the single-target relay path has no fork-aggregator final_forwarded to dedup it), so no second final reaches the UAC. On a final (≥200) — UAS already answered — it is a no-op returning False (a proxy cannot retract a 2xx); the script branches on the bool (log + reply.relay(), best-effort). Takes precedence over relay(). Unit-tested (script::api::reply::tests decision logic + proxy::session::tests flag) + SDK-mirrored/tested (reply.reject, sdk/tests/test_reply_reject.py). End-to-end SIPp-validated (sipp/reject_{uac,uas}.xml + reject_proxy.py): caller gets 100→503 Media Authorization Failed (To-tag added, 183 suppressed), UAS gets the CANCEL on the INVITE's Via branch (RFC 3261 §9.1) and its 487 is ACKed and absorbed — both endpoints 1 Successful / 0 Failed / 0 Retrans / 0 Unexpected, siphon 0 WARN/ERROR. SIPp validation also surfaced + fixed a pre-existing loop: a non-compliant ACK (fresh branch instead of the INVITE's, §17.1.1.3) carrying the 503's To-tag + an R-URI pointing at the proxy was matched by by_dialog_key and relayed to the proxy's own address in handle_ack_via_session, stacking a Via per hop until the datagram exceeded the 8192-byte UDP buffer (truncated → parse-error drop). Two guards: (1) reject now drops the dead by_dialog_key entry (ProxySessionStore::remove_dialog_key — a rejected INVITE forms no dialog), and (2) handle_ack_via_session reuses the existing is_own_address loop check to silently drop an ACK whose resolved next-hop is one of our own listeners (RFC 3261 §16.3). |
| Session timers (RFC 4028) | Implemented | session_timer: |
One session timer per dialog: the callee's negotiated from its 2xx (§7.2), the caller's from the caller's INVITE (§9 Table 2). refresher (uac/uas/b2bua, also call.session_timer(refresher=), which runs a timer without the block) is the preference where the negotiation leaves siphon the choice. siphon refreshes the dialogs it is the refresher of and ends a call whose refresher let a session run out, a third of the interval (at most 32 s) early (§10); a refresh answered 408/481 or unanswered for 64·T1 ends the call, a 422 retries at the larger Min-SE. A party that supports the extension and asks for less than siphon's min_se is refused 422 with siphon's minimum in Min-SE (§9) instead of being taken at an interval siphon's refreshes then raise: an initial INVITE once the script has decided, before the call is dialled, handed over or answered, and a re-INVITE or UPDATE refresh from either party on its own dialog, never relayed. A refresh is a request on that leg's own dialog (its Call-ID, tags, route set, CSeq and remote target, siphon's Contact; toward the callee with its INVITE's Supported/Require/Proxy-Require), never a copy of the other leg's request, offering the session description in force there (the offer last accepted, the answer last sent, early media SDP when the 2xx had none), so a held call stays held and an unchanged session keeps its o= version (RFC 3264 §8). With none in force: a bodyless UPDATE where the peer allows UPDATE, else a bodyless re-INVITE whose 2xx offer is relayed to the other party, whose answer goes in the ACK (a refusal or no answer ends the call, the ACK rejecting every stream ahead of the BYE). The same rules run on calls siphon answers itself (call.answer(), a handover's answer, the control plane's: §9 from the caller's INVITE, a call.session_timer() set before call.answer() included), on calls siphon places (the INVITE asks for the configured timer or b2bua.originate(session_timer={...}), the callee's 2xx names the refresher) and on bridge legs (every re-INVITE siphon sends on a dialog carries its timer or asks for siphon's, and its 2xx sets the timer and the session in force). An SDP siphon sent as given (a script's or media engine's answer, an originate's offer) keeps its own o= session id, so a refresh offers it unchanged. End-to-end SIPp scenarios run on every PR: b2bua-session-timer, and b2bua-session-timer-floor at Session-Expires 90 / Min-SE 90 on its own instance, where siphon refreshes the callee's dialog at half the interval and ends the call on both legs when the caller, the refresher of its own dialog, stops refreshing, each within a time window both SIPp parties assert. Dispatcher-driven tests in dispatcher::session_refresh_tests, dispatcher::session_timer_tests, dispatcher::session_timer_legs_tests and dispatcher::session_timer_min_se_tests. |
| UPDATE before the answer (RFC 3311 §5.2) | Implemented | Core | An UPDATE from the caller of a B2BUA call whose callee has not answered: no offer 200; an offer while the INVITE's offer is unanswered 500 with Retry-After 0-10 s; an offer after the caller acknowledged its answer in a reliable provisional 504. Nothing reaches the media engine or the callee (dispatcher::early_update_tests, on an in-process native engine). Not implemented: relaying such an offer to the callee's early dialog. |
| PRACK (RFC 3262) | Implemented | Core | Reliable provisional responses. The B2BUA owns 100rel per leg under every preset: it auto-PRACKs a reliable-provisional B-leg, and the callee's Require: 100rel/RSeq never cross. Toward the caller a provisional is reliable on siphon's own RSeq (one more per provisional on the caller's dialog) when the caller requires 100rel, or supports it and the callee sent that provisional reliably or it carries SDP; retransmitted on T1 doubling until the caller's PRACK; one reliable provisional at a time. siphon answers the PRACK on RAck + dialog (200, again for a retransmission or after the final; 481 unmatched); a 2xx is held until the PRACK of a reliable provisional with SDP, and only once sent is it retransmitted and under the 64T1 no-ACK sweep and the §15 BYE hold; a teardown or CANCEL while it is held gives the caller 487 and the callee a BYE; retransmits stop on the final; no PRACK in 64T1 refuses the caller 500 through the teardown claim and CANCELs, or BYEs, the callee. call.progress()/call.answer() follow the same rules. Offer/answer in PRACK crosses the call (RFC 3262 §5): siphon's PRACK to the callee waits for the caller's and carries its SDP (the answer to an early offer, or a new offer whose answer returns in the 200 to the caller's PRACK, with a callee 2xx waiting behind that 200), through the media engine when anchored; a missing or refused answer, or 64T1 without one, ends the call through the teardown claim. Retransmissions of a reliable provisional siphon already has are discarded (§4). A PRACK still waiting for the caller's when the caller gets its 2xx goes to the callee then, to the answering branch but not to one that failed or was CANCELled. What a PRACK exchange agrees is the session description in force a later session refresh offers (the answer in siphon's PRACK as it goes, an offer once the callee answers it), and so is the SDP siphon sends the caller in a reliable provisional (an answer as it goes, siphon's offer to an INVITE without SDP once the caller's PRACK answers it), with the o= session id and version it went with as the caller's dialog's, so a refresh toward the caller offers it unchanged (RFC 3264 §8). An offer in a PRACK that arrives after the caller's 2xx, when siphon's PRACK already went with it, reaches the callee in an UPDATE (RFC 3311) and its answer returns in the 200 to that PRACK; a refusal leaves the session and refuses the PRACK the same way (488 for 488/606, 491 and 504 as they are, 500 with Retry-After for 500, 500 otherwise), a 481, a 408 or 64T1 without a response ends the call (§5.3), and a callee without UPDATE in Allow gets none while the caller gets 488. A new PRACK with an offer is a new offer, not a retransmission: 500 with Retry-After while an earlier one is still out, otherwise an UPDATE of its own; the caller's offer becomes its own SDP only once the callee accepts it. The target of a leg replacement (a siphon-terminated REFER, replace_peer) is PRACKed at once on its own early dialog whatever the caller supports, since its provisionals are never relayed to a caller that could PRACK them. A PRACK, and a branch's recorded failure, follow the leg by the Via branch of its INVITE, not by its position on the call, so a leg ahead of it taken off the call in between does not misdirect them (dispatcher::replacement_prack_tests, b2bua::actor::store::by_branch). The rest of the B-leg response path follows the Via branch the same way (the answer's claim and its ACK, a refused answer's release, a 401/407 or 422 retry, an LCR carrier's failure, the settling of a relayed re-INVITE, UPDATE, REFER, NOTIFY or INFO and of a request siphon sent on a leg, and the ACK held for a delayed offer): dispatcher::b_leg_position_tests runs each handler with the call read before a leg ahead was taken off it, and delayed_offer_ack_tests::a_held_ack_follows_its_leg_when_a_leg_ahead_of_it_is_taken_off_the_call does the same for the held ACK. dispatcher::a_leg_reliable_provisional_tests, dispatcher::prack_offer_answer_tests, b2bua::actor::reliable, b2bua::actor::prack_bridge, dispatcher::b2bua::late_prack_offer. End-to-end SIPp scenarios: b2bua-reliable-prov (non-100rel A-leg, against a dedicated sip-trunk-edge@2026 instance whose response default is Copy) runs on every PR; b2bua-a-leg-reliable-prov (callers that require and that support 100rel, the early offer answered in a PRACK, an offer in a PRACK, and a late PRACK offer carried in an UPDATE, against an ims-intra-trust-domain@2026 instance whose response policy copies RSeq; CI job sipp-b2bua-a-leg-reliable-prov) runs on every PR. |
| E.164 number normalization (identity headers) | Implemented | numbering: / number_policies: / request.rewrite_identities() / call.dial(number_policy=…) |
One call reformats every dialable identity userpart (From, To, P-Asserted-Identity, P-Preferred-Identity, Request-URI, opt-in Referred-By/Remote-Party-ID) into e164/plain/international/national under a home numbering plan + named versioned presets. Display names, tags, hosts, non-numbers and preserve_users service/emergency codes untouched; a national form of a foreign number falls back to the international access form. numbers.parse(raw, home=) exposes .e164/.plain/.international/.national/.cc/.nsn/.format(). B2BUA number_policy= (or b2bua.default_number_policy) normalizes the A-leg identities that flow to the B-leg plus the dial/fork target; an LCR route's number_policy (else the same default) does the same for its carrier's R-URI and identities. Opt-in diversion: extends to Diversion (RFC 5806) / History-Info (RFC 7044) with structured per-entry rewrites preserving index/reason/embedded escaped cause/ordering and privacy-restricted entries (respect_privacy). Rust unit + KAV-vector + script-engine end-to-end tests; SDK-mirrored + tested (sdk/tests/test_numbers.py). |
| B2BUA header policies | Implemented | header_policies: / b2bua.default_header_policy / call.dial(header_policy=…) |
Which headers cross the A-leg/B-leg trust boundary, as named versioned policies instead of per-call strip/copy logic. Four built-in presets ship (transparent-b2bua@2026 — the default, behaviour-equivalent to pre-policy siphon; ims-intra-trust-domain@2026; ims-trust-domain-boundary@2026; sip-trunk-edge@2026), and header_policies: defines operator-owned ones in the same namespace — either extends: a built-in with copy/strip/rewrite/translate deltas (a direction left out is inherited verbatim; an extends: with no rules is a stable local alias), or a full declaration with an explicit per-direction default:. Header names are exact and case-insensitive with a trailing * prefix match; within a direction an exact name beats a prefix and a longer prefix beats a shorter one, so strip: ["X-*"] + copy: ["X-Account-Ref"] is expressible in one block. Per-call copy=/strip=/translate= deltas still layer on top and still win (exact names only), and a header the script set or removed on the A-leg INVITE (call.set_header/remove_header/remove_headers_matching) goes out as the script left it over both: no preset or delta strips, rewrites or translates it, framework-managed headers excepted (script_header_precedence_tests). Likewise a header a script set or removed on a relayed response (reply.set_header/remove_header/remove_headers_matching in @b2bua.on_answer/@b2bua.on_early_media) reaches the caller as the script left it, over the response policy and siphon's own Supported/Allow (reply_script_header_tests). Policies are resolved and validated at config load: an unknown op token, a rule aimed at a framework-managed header (Via/Call-ID/CSeq/Max-Forwards/Content-Length/From/To/Contact/Record-Route/Route), an unversioned name, a built-in name collision, or a default_header_policy naming nothing all refuse to start — the default previously warned and fell back to the most permissive preset, so a typo opened the boundary it was meant to close. The B-leg INVITE's Supported and Allow are siphon's own whatever the policy, since siphon is that leg's UAC (RFC 3261 §20.37/§20.5): Allow is siphon's method set, Supported is replaces plus the caller's 100rel/timer, and the caller's end-to-end tags only under a policy that copies what each negotiates with both ways: precondition needs Supported+Require (every built-in except transparent-b2bua@2026), histinfo needs History-Info (transparent, intra-trust), resource-priority needs Resource-Priority out and Accept-Resource-Priority back (every built-in except the trust boundary). Other caller tags (outbound, path, gruu, …) no longer go out as siphon's claim; a script set_header/remove_header of either header goes out as written. Responses relayed to the caller follow the same rule mirrored, under every preset: siphon's Allow, the callee's 100rel/timer/relayed end-to-end tags plus replaces (the default preset strips the callee's Supported first, so its responses stay Supported: replaces) (b_leg_capability_tests). Rust unit tests (resolution, precedence ordering, one per rejection) + config parse tests + integration test driving config → registry → applied B-leg INVITE. |
Transports¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| TCP | Production | listen.tcp |
AS-facing; RFC 3261 §18.3 stream framing with Content-Length extraction; outbound distributor falls back to the ConnectionPool when an OutboundMessage arrives without a matching inbound connection (covers UAC fire-and-forget paths like in-dialog NOTIFY from subscribe_state.notify() whose Route header points at a destination with no live inbound socket — previously the message was built but silently dropped at the connection-map lookup). Wedge-hardened (all stream transports): the per-listener outbound distributor routes with a non-blocking try_send instead of send().await. A single non-reading peer (toll-fraud scanner that never ACKs its 401s, or a stream peer whose far end stalls) fills its bounded per-connection channel; an awaiting send parked there while holding the connection_map shard read guard, stalling outbound for every connection (head-of-line) and blocking the accept loop's insert on the same shard — accept stops, the backlog fills, the engine wedges (no logs) until restart. try_send keeps the guard only for the synchronous send and sheds a backed-up peer. Reproduced + regression-guarded black-box on a real container at --cpus 0.5 by scripts/wedge_test.sh (run-tests.sh --wedge) — probe times out pre-fix, answered post-fix. Outbound ConnectionPool establishment hardened: the connect is bounded by a fail-fast timeout (TCP_CONNECT_TIMEOUT, 5 s) so a doomed ESP-over-TCP send to a UE whose IPsec SA was just torn down (no SYN-ACK, no RST) can no longer block the PyExecutor worker indefinitely and trip the script-executor watchdog → process abort; and concurrent first-sends to the same destination coalesce onto one connection under a per-destination lock, so the fixed protected source port (pcscf_port_c) cannot hit EADDRNOTAVAIL/EADDRINUSE on a second bind/connect of the same 4-tuple. Regression-guarded by connect_fails_fast_to_blackhole and concurrent_sends_coalesce_onto_one_connection in transport::pool. A connection's cleanup only evicts its own pool entry, so the reader of a stalled connection that was already replaced no longer drops the replacement when it ends (an_old_connection_closing_does_not_evict_its_replacement). Inbound flow reuse (RFC 5923 / RFC 5626 §5.3): accepted TCP connections (dedicated listener and the SIP half of the TCP+WS mux) are registered in the StreamConnections registry for their lifetime and evicted on close, as TLS/WS are, so Flow.is_alive and the subscribe_state received-flow NOTIFY work for a peer reachable only over the connection it opened. Opt-in: a URI relay over TCP still dials through the pool. The registry is keyed by address + transport, so a TCP entry is never picked for a TLS send. Tested end to end on loopback (transport::tcp::tests, incl. the drain-to-baseline leak gate) and in transport::mux::tests. |
| TLS | Production | listen.tls |
Subscriber-facing, TLS 1.3 validated; RFC 3261 §18.3 stream framing; outbound distributor wedge-hardened with non-blocking try_send (see TCP) |
| TLS 1.3 | Production | tls.method: TLSv1_3 |
tls.method is the minimum version and is enforced on both the inbound acceptor and the outbound connection pool: TLSv1_3 negotiates 1.3 only and refuses a TLS 1.2 peer in either direction. Previously the value was parsed into config and never read — the acceptor was built from a bare rustls::ServerConfig::builder() — so a config claiming a 1.3 floor still handshook with TLS 1.2 clients. Unit-tested against real handshakes with version-pinned peers. |
| TLS 1.2 | Implemented | tls.method: TLSv1_2 |
The default, and a floor rather than a pin: a 1.3-capable peer still negotiates 1.3. Matches what siphon has always served, so an unset or TLSv1_2 config is unchanged. TLS 1.0/1.1 and SSL are rejected at config load (RFC 8996). |
| mTLS — inbound (verify client cert) | Implemented | tls.verify_client: true, tls.client_ca |
Client certificate required and verified against the tls.client_ca PEM bundle; applies to listen.tls and listen.wss (shared TLS block). Fails closed at startup if verify_client is set without client_ca (previously verify_client was silently ignored on the SIP listener — read only by the X1 LI interface). TLS handshake bounded by a 10 s timeout (half-open-handshake / slowloris defense). |
| mTLS — outbound (present client cert) | Implemented | tls.client_certificate, tls.client_private_key |
Siphon presents this client certificate on outbound TLS connections whose peer requests one — for upstream SIP trunks requiring client-certificate / mutual TLS (e.g. Teams Direct Routing), which previously aborted the handshake with CertificateUnknown because the outbound pool presented no client cert. Both fields must be set together or neither; a one-sided setting or an unreadable/unparseable file is a hard startup error (fail closed). Server-certificate verification is unchanged (permissive). Unit-tested against a real handshake (mandatory-mTLS server built with rcgen): matching identity succeeds, no identity is rejected. |
| Outbound TLS SNI (RFC 6066) | Implemented | automatic | Outbound TLS handshakes present the resolved target hostname as SNI / certificate name instead of the destination IP literal (rustls sends no SNI for an IP). The hostname flows from the resolved SIP URI through relay, fork, and the gateway TLS health probe into the connection pool; bare-IP next hops send no SNI (unchanged). |
| Inbound TLS SNI — per-domain certificates (RFC 6066) | Implemented | tls.certificates[] |
Serves a different certificate per server name on one listen.tls / listen.wss socket (the OpenSIPS tls_mgm / Kamailio tls_domain equivalent), so multi-domain edges stop needing one SAN certificate that couples every domain to a single ACME renewal. Exact names and single-label wildcards (*.example.com, RFC 6125 §6.4.3 — not the bare domain, not a deeper subdomain), matched case-insensitively (RFC 4343), exact winning over a wildcard covering it. tls.certificate/tls.private_key stay the default for a client sending no SNI (any IP-literal peer) or an unmatched name — selection never aborts a handshake, so a config with no certificates: behaves exactly as before. Every pair hot-reloads independently on its own renewal. Fails closed at startup on a duplicate server name, an entry with no server_names, a malformed wildcard, or an unreadable/mismatched pair (the path is named in the error). verify_client/client_ca remain listener-wide. Unit-tested plus an end-to-end handshake assertion that the certificate a client receives is the one its SNI selected. |
| UDP | Production | listen.udp |
One SO_REUSEPORT socket per worker. Every destination is routed to one worker's outbound channel (transport::udp::UdpOutbound), so messages to one peer leave in the order they were sent, whatever call site sent them; before, all workers drained one channel and two back-to-back messages to a peer could arrive inverted (a BYE before its ACK, a 487 before the 200 to its CANCEL). Unit-tested in transport::udp::tests (400 separately enqueued messages to one peer arrive in order, destination routing, a failed worker's channel still drained). |
| WebSocket (WS) | Implemented | listen.ws |
RFC 7118, browser/WebRTC clients; outbound distributor wedge-hardened with non-blocking try_send (see TCP). MT routing (INVITE → WS-registered UE) works via RFC 5626 §5.3 connection reuse: every binding captures its inbound flow (no flow_token= needed), registrar.lookup() returns it as contact.flow, and request.fork(contacts) / request.relay(flow=) / call.fork(contacts) / call.dial(flow=) route over the captured connection on both proxy and B2BUA. Connections register in a unified cross-transport StreamConnections registry (also backs Flow.is_alive); send_to_target has a WS/WSS arm that reuses the connection and drops (no caller-echo) on miss. |
| Secure WebSocket (WSS) | Implemented | listen.wss |
Outbound distributor wedge-hardened with non-blocking try_send (see TCP). MT routing via connection reuse — same flow-based path as WS (see the WS row). |
| Shared-socket SIP + WebSocket mux | Implemented | same address in listen.tls + listen.wss, or listen.tcp + listen.ws |
One listening socket carries raw SIP and the RFC 7118 WebSocket upgrade, classified per connection from its first line (SIP/2.0, RFC 3261 §7.1/§7.2, versus HTTP/1.1, RFC 6455 §4.1 — disjoint grammars, so the split is exact). One port, one firewall pinhole and one certificate for a browser UE on WSS and a SIP trunk on TLS. Each connection is then handled exactly as on a dedicated listener and stamped with the transport it speaks, so Via/Contact, flow capture, MT routing and outbound distribution are unchanged; cost is one classification per connection, nothing per message. A silent peer (connection reuse, RFC 5923) falls back to raw SIP after 2 s; a probe that is neither protocol is dropped and counted as malformed. Only tcp+ws and tls+wss may share a socket — any other pairing on one address (plaintext with TLS above all) is a startup error, replacing the previous silent SO_REUSEPORT split where the kernel divided connections between two listeners. Unit-tested (first-line classifier, prefix replay, seeded framing) plus end-to-end over real loopback sockets for both pairings, including responses on each half's own framing. |
| SCTP | Implemented | listen.sctp |
RFC 4168, IMS inter-node; outbound distributor wedge-hardened with non-blocking try_send (see TCP). Inbound only: the distributor writes to associations the listener accepted, and the outbound connection pool opens no SCTP — there is no outbound SCTP dialer, so siphon cannot originate towards an SCTP peer it has not heard from. The egress channel is therefore created only when a listen.sctp entry will read it; with none (or a binary built without the sctp feature, where the block is warned about and ignored) the router refuses a Transport::Sctp send at ERROR naming the destination and counts it on siphon_outbound_unserved_total{transport}, instead of queueing it unbounded behind a reader that does not exist. Unit-tested in transport::tests::{no_sctp_listener_means_no_sctp_egress_channel,outbound_router_refuses_sctp_when_nothing_serves_it} |
| Per-socket advertised address | Production | listen.<transport>[].advertise |
Host and optional port (host, host:port, [v6]:port), parsed at config load; a malformed value is a startup error. A configured port replaces the bound port in Via, Record-Route, B2BUA Contact, the answered-OPTIONS Contact and UAC-originated Via/Contact, and joins the Route self-identity; bind and send are unchanged (unit and end-to-end dispatcher tests). At startup, siphon warns for each TLS or WSS listener that advertises an IP literal without a matching iPAddress subjectAltName in tls.certificate. A peer opening a new TLS connection to that host fails certificate validation and never sends the ACK or in-dialog request. The check is unit-tested against rcgen certificates. |
| Global advertised address | Implemented | advertised_address: |
Fallback for 0.0.0.0 binds. Host only, same grammar as advertise; a port is refused at startup (one port cannot fit every transport). |
| DSCP/ToS marking | Implemented | listen.dscp |
RFC 4594 signaling QoS; default CS3 (24); per-listener override |
| PROXY protocol (client address from a front) | Implemented | listen.{tcp,tls,ws,wss}[].proxy_protocol.from |
HAProxy PROXY protocol v1 (text line) and v2 (binary), read at accept on any stream listener and on a shared tcp+ws / tls+wss socket — before the first-line SIP sniff and before the TLS handshake, so a front may terminate the UE's TLS and re-encrypt to siphon with no cleartext SIP on any hop. The header's source address then replaces the front's for every consumer that keys on the source: failed_auth_ban, from_gateway() / source_ip_in(), NAT return-routing (received=/rport=), media.received_from (which otherwise gated RTP ingress to the front, so no media was accepted at all), capture and the CDR. L4 source preservation (externalTrafficPolicy: Local, host networking, DSR) is the fix only for a front that forwards packets; it cannot work for one that terminates the connection, which is the point of using one to terminate TLS. from is a mandatory CIDR allowlist with no default and no inheritance from security.trusted_cidrs — that list means "exempt from abuse controls" to all four of its consumers and is where monitoring boxes and trunks are listed, so inheriting it would hand source-address forgery rights to hosts named there for an unrelated reason. Refused at config load: the option on a udp listener (a UDP reply goes to the peer address, so a substituted source would send every answer past the front), an empty from, a non-CIDR entry. On an enabled listener a sender outside from and a connection opening without a header are both dropped and never attributed to the front, and neither credits the auto-ban store (a second front nobody added to the list is likelier than abuse); v2 LOCAL / v1 UNKNOWN — a front's own health checks — are consumed and the socket's peer address stands. A PROXY header on a listener with the option off is held apart from Sniff::Garbage and dropped with a log naming the listener, instead of scoring a strong auto-ban signal: siphon banning its own load balancer. The v2 PP2_TYPE_SSL TLV (the client's TLS session at the front) is carried beside the hop, never over it: request.transport keeps naming the transport siphon accepted, because that decides the connection map, the pool and the Via/Contact token, while request.client_transport and request.client_is_secure answer for the client. client_is_secure answers for the client's effective hop, so a UE connecting straight to a tls listener reports true as well — reading a missing client_transport as "insecure" would lock out every direct-TLS UE. Contact.client_transport persists it with the binding and the CDR records the client's transport; mirrored in the SDK. Unit-tested (transport::proxy_protocol — v1/v2 parse, headers split across reads, LOCAL/UNKNOWN, the allowlist, replay of the over-read past the header; transport::stream first-line classification, including that a PROXY header is held apart from garbage and that prefix-matching does not eat a CRLF keepalive; config::tests for each refusal), and each of those behaviours is mutation-proven rather than merely passing. |
Registrar¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Redis backend | Production | registrar.backend: redis |
Persistent across restarts |
| Memory backend | Implemented | registrar.backend: memory |
Ephemeral |
| PostgreSQL backend | Implemented | registrar.backend: postgres |
|
| Python custom backend | Implemented | registrar.backend: python |
|
| Expires control (default/min/max) | Production | registrar.{default,min,max}_expires |
An Expires below min_expires is answered 423 Interval Too Brief with Min-Expires by registrar.save(), which returns False and stores nothing |
| Max contacts per AoR | Production | registrar.max_contacts |
A new binding past it is answered 503 Service Unavailable with Retry-After (seconds until the soonest held binding expires) by registrar.save(), which returns False. The REGISTER is applied atomically: a refusal stores none of its Contacts, and force=True clears only once every Contact is accepted. Refusals count in siphon_registrar_refusals_total{reason}, not as script errors |
| Bind AoR to authenticated user | Implemented | registrar.enforce_auth_aor_match |
Rejects (403) a REGISTER whose AoR (To-URI user) ≠ the authenticated digest user — anti account-takeover / forced-deregister. Checked before the force-clear so a spoofed AoR can't first wipe the victim's bindings. Default off (IMS deployments authorize via the implicit registration set, where the public identity ≠ private auth identity). |
| Redis TTL slack | Production | registrar.redis.ttl_slack_secs |
Race condition buffer |
| GRUU (RFC 5627) | Implemented | ||
| Service-Route (RFC 3608) | Production | Via registrar.set_service_routes() / service_route() |
|
| Registration state change hooks | Production | @registrar.on_change |
Callbacks on insert/delete/expire |
| Liveness — flow-failure dereg (RFC 5626 §4.2.2) | Implemented | registrar.liveness.enabled |
TCP/TLS/WS/WSS connection close (peer FIN/RST, read error, idle timeout, or CRLF-keepalive failure) deregisters the bindings that arrived on that connection. Transport notifies the registrar over a close channel → Registrar::unregister_flow(connection_id), which uses a ConnectionId → AoR reverse index (connection_index) to drop only the affected bindings (O(bindings-for-that-connection), scanner-churn-safe) and emit Deregistered. Default off. |
| Liveness — IPsec idle dereg (UDP + TCP/TLS/WS) | Implemented | registrar.liveness.{enabled,keepalive_interval_secs,idle_multiplier,probe_timeout_ms} |
Detects a dead UE on the production Gm without a SIP de-REGISTER, on any SIP transport — the XFRM SA use-time is the liveness signal, so a TCP+IPsec registration whose UE silently dies (radio loss, no FIN/RST) is reaped on the same ~idle_multiplier × keepalive_interval window as a UDP UE, rather than waiting for the CRLF-keepalive timeout (minutes). The 30 s sweep polls kernel XFRM SA inbound use-time (one XFRM_MSG_GETSA netlink dump — no per-packet hot-path cost; the UE's RFC 6223 keepalive keeps the SA warm); eligibility is by SA match (UE IP), which naturally excludes non-IPsec bindings. A suspect binding is probed with one OPTIONS over its actual transport (stream probes ride the captured inbound connection); no answer → deregister. SA teardown (sweep_expired) also drops the matching binding. EPC-independent backstop for SMF crash / PCRF-no-Rx-ASR / silent radio loss. Idle-probe path needs kernel XFRM + lab validation. Default off. |
| Liveness — network dereg cascade | Implemented | registrar.liveness.dereg_mode: network_dereg\|local_only |
Removing a binding emits Deregistered → @registrar.on_change (the authoritative/S-CSCF path; the script sends the terminated reg-event NOTIFY). For a P-CSCF cache binding (carries a flow_token) under network_dereg, also synthesizes a de-REGISTER (Expires: 0) on the UE's behalf routed via the stored Service-Route so the registrar of record clears it too. local_only skips the upstream REGISTER. Network-dereg routing needs a split P-CSCF/S-CSCF lab to validate end-to-end. |
| Outbound registration (registrant) | Production | registrant: |
UAC REGISTER to upstream trunks. A registered TLS or TCP trunk whose own outbound pool connection to the registrar (exact address, port and transport) is gone re-registers on the next 5 s tick instead of at the refresh timer; an inbound connection from the registrar's host never counts as that connection. A trunk's connection is exempt from the pool's 30 s TCP idle timeout, so silence on an idle trunk is not mistaken for a dead one (it was re-registering about every 35 s whatever the interval said). Detection is unaffected: a registrar that closes, errors or stops answering its socket keepalive (SO_KEEPALIVE 60 s + 3 × 10 s, TCP_USER_TIMEOUT 90 s) loses the connection and the trunk re-registers within 5 s of that. There is no outbound SCTP connect path, so transport: sctp (and ws/wss, and a typo) is refused at config load naming the entry and what to use instead, and registration.add(transport=) raises ValueError; all four used to parse and then register over UDP, as did a mis-cased "TLS". Accepted set is udp/tcp/tls, case-insensitive, from one parser (config::parse_outbound_transport) shared with the gateway destinations and the database/HTTP row checks. Unit-tested in registrant::liveness_tests and transport::pool::tests::{has_live_*,*idle_exempt*}. |
| Outbound registration — database / HTTP source | Implemented | registrant.backend: database | http |
The registering trunks are read from a source the controller owns and reconciled on refresh_secs: adds, re-registers a changed row (same Call-ID, continuing CSeq — a refresh, not a new registration, RFC 3261 §10.2.4), and de-registers (Expires: 0, §10.2.2) one the source dropped or disabled. An unchanged row is untouched, so polling does not churn the estate; an unreadable source changes nothing rather than tearing it down. The source owns only entries it created — registrant.entries and registration.add() are never reconciled away. Secret is a password or a pre-computed ha1 (RFC 7616 §3.4.3 — bound to its hash, and a mismatch is reported rather than answered wrongly). Versioned JSON contract typed in siphon_sdk.registrants, documented at docs/reference/registrant-api.md. Pure-function reconcile diff plus source-ownership and identity-preservation tests; not yet validated end-to-end against a live provisioning source. A failed poll backs off from 250 ms capped at refresh_secs instead of sleeping the whole interval, and a streak escalates warn → error from its second poll, throttled to a line a minute. siphon_registrant_source_last_success_timestamp_seconds / _failures_total make it measurable, tracked separately from the gateway pair because they fail separately. This loop is spawned after the listeners bind and keeps that ordering: nothing inbound depends on a trunk registration, so the worst a slow first read costs is a late REGISTER. |
| IMS UE registration (soft-UE, AKA + IPsec sec-agree) | Implemented | registration.add(auth="aka", k=, opc=, ipsec=True, ue_port_c=, ue_port_s=) or YAML registrant.entries[].{auth: aka, aka:, ipsec:} |
siphon registers INTO an IMS core as a handset: IMS-AKAv1-MD5 (RFC 3310 — RES is the binary digest password) over IPsec sec-agree (3GPP TS 33.203). Milenage f1*/f5*/AUTS re-sync (TS 35.208 Test Set 1 vectors). Initial REGISTER offers Security-Client (UE SPIs/ports + Require: sec-agree); the 401 records Security-Server; the protected re-REGISTER echoes Security-Verify and egresses from the UE protected client port over the four UE-side SAs (create_ue_sa_pair — same netlink + CK/IK derivation as the P-CSCF, only the four XFRM policy directions mirror via SaRole); the protected 200 OK tightens the SA hard-lifetime to the granted Expires + Timer-F grace. Service-Route / P-Associated-URI captured for MO routing; AUTS re-sync is a fallback (a fresh stateless UE never emits it). Message construction unit-tested against 3GPP/RFC vectors; the kernel SA install is root-gated. NOT yet validated end-to-end against a live P-CSCF. Example: examples/ims_ue_b2bua.{py,yaml}. |
| IMS UE B2BUA bridge (plain SIP ↔ IMS) | Implemented | examples/ims_ue_b2bua.py, call.dial(flow=, route=), registration.flow()/service_route() |
Bidirectional B2BUA over the soft-UE registration. MT (IMS→tester): the protected-port A-leg bridges to a plain-SIP tester; A-leg responses egress back over the SA via inbound-flow pinning. MO (tester→IMS): dials the B-leg over the UE→P-CSCF SA flow (registration.flow(impu, ue_ip) → call.dial(flow=), sourced from the UE protected client port), carrying the captured Service-Route (registration.service_route(impu) → call.dial(route=)) and asserting the IMPU via P-Preferred-Identity (intra-trust preset preserves P-*). Direction detected by call.source_ip == pcscf. SDK-tested both directions (sdk/tests/test_ims_ue_b2bua.py); needs live-core + root validation. |
| Proxy-side binding cache | Implemented | registrar.save_proxy(request, reply) |
P-CSCF caches what S-CSCF granted; reads Expires from reply (not request), bypasses local max_expires cap, +32 s Timer F grace, no auto-200 OK (proxy relays upstream's response) |
| Path-token MT routing (RFC 3327 / TS 24.229 §5.2.7.2) | Implemented | request.add_pcscf_path(token), registrar.save(flow_token=)/save_proxy(flow_token=), registrar.lookup_by_token(token), request.relay(flow=binding.flow), ipsec.path_host |
P-CSCF mints opaque token, embeds in Path userpart; binding stores token + captured inbound flow (source addr, listener local addr, accepted-connection id); MT routing bypasses DNS resolution and egresses from the same listener that received the REGISTER. UDP flow survives restart; TCP/TLS/WS/WSS bound to accepting instance lifetime. Via on flow-relay derives from flow.local_addr so IPSec port pairs are preserved (TS 33.203 §7.4). |
| AS-side contact capture (TS 24.229 §5.4.2.1.2) | Implemented | registrar.save_as_contact(aor, reply), Contact.params, Contact.kind |
S-CSCF script caches the AS's Contact: URI and RFC 3840 feature tags (+g.3gpp.smsip, +g.3gpp.icsi-ref, …) from the 3PR 200 OK. AS contacts are excluded from registrar.lookup() (routing-side never picks them as MT targets), from registrar.reginfo_xml(...) by default (RFC 3680 §5.2 — an AS that answered a 3PR has not registered a contact for the user; opt in with include_as_contacts=True for a watcher that wants them, where they render as <unknown-param> children per §5.3.2), and cascade-clear when the last UE binding deregs/expires. The UE-facing capability surface is RFC 6809 Feature-Caps on the REGISTER 200 OK. |
| Contact-header parameter passthrough (RFC 3840) | Implemented | Contact.params |
Every non-typed Contact-header parameter (anything outside tag/q/expires/+sip.instance/reg-id) round-trips through save → backend persistence → lookup → reg-event NOTIFY. Lowercased at parse time per RFC 3261 §19.1; values preserved verbatim. |
Authentication¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Digest auth — 401 (UAS) | Production | auth.require_digest() |
REGISTER challenges. The offered algorithm set is auth.algorithms (default [MD5, SHA-256, SHA-512-256], unchanged): one WWW-Authenticate per entry in the configured order, sharing one nonce (RFC 7616 §3.7). Narrowable to a single algorithm for a client population that abandons a registration on any multi-challenge 401 — the SHOULD assumes a client picks one from the list, and some do not. Verification is independent of the offer and still accepts any algorithm siphon can compute. Unknown names, an empty list and AKAv1-MD5 are refused at config load rather than skipped. Unit-tested in script::api::auth::tests (default set and order, single-entry set, configured order on the wire, one nonce across the set, re-challenge replaces rather than stacks) and config::tests (the three refusals, each naming the offending value). |
| Digest auth — 407 (proxy) | Production | auth.require_proxy_digest() |
INVITE challenges; same auth.algorithms set, emitted as Proxy-Authenticate. |
| Digest auth — B2BUA A-leg | Implemented | auth.require_proxy_digest(call, …) |
Challenge the caller from @b2bua.on_invite. With a @b2bua.* handler registered the INVITE never reaches @proxy.on_request, so the digest helpers take the Call too; the 407 is armed as the call's deferred reject, so no B-leg is dialled, and the caller's hop-by-hop Proxy-Authorization is stripped before it could reach one. Verified username on call.auth_user and on the CDR. Gated by scripts/run-tests.sh --b2bua-invite-auth. |
| Multi-algorithm challenge (RFC 7616 §3.7) | Implemented | One WWW-Authenticate/Proxy-Authenticate per algorithm (MD5, SHA-256, SHA-512-256) on a single 401/407, so RFC 2617 and RFC 7616 clients both negotiate. Wire shape validated against the Wireshark dissector. |
|
| HTTP backend (HA1 lookup) | Production | auth.backend: http |
REST credential lookup; optional per-username TTL cache (auth.http.cache_ttl_secs) flattens registration storms so repeat REGISTERs skip the blocking fetch |
| Static users backend | Implemented | auth.backend: static |
Inline config credentials |
| Diameter Cx backend (HSS) | Production | auth.backend: diameter_cx |
3GPP TS 29.228 |
| AKA / AKAv1-MD5 (HSS-backed) | Production | auth.require_ims_digest() |
3GPP TS 33.203 via Cx MAR/MAA |
| AKA / AKAv1-MD5 (local Milenage) | Implemented | auth.aka_credentials |
3GPP TS 35.206 — local key derivation without HSS. The 401 carries ck=/ik= like the HSS path, so a P-CSCF in front sets up IPsec from it; exercised end-to-end over real XFRM SAs in the sipp-ipsec harness (sipp_ipsec UE) |
| SHA-256 digest (RFC 7616) | Implemented | ||
| Anti-spoofing (from=auth check) | Production | Script logic | auth_user == from_uri.user (caller-ID/From); the registrar-side AoR/To equivalent is registrar.enforce_auth_aor_match |
| Digest nonce replay protection (RFC 7616 §3.3) | Implemented | auth.nonce_secret, auth.nonce_ttl_secs |
Nonces are timestamp-bound ({unix_secs:016x}.{tag}) and rejected once older than the TTL (default 3600 s), bounding captured-Authorization replay from "forever" to the window. Cross-instance safe with no shared state (correct behind round-robin DNS where a re-REGISTER may land on a different node). Optional shared nonce_secret adds HMAC-SHA256 integrity so a node rejects nonces the cluster never issued — must be identical on every instance behind the domain. Applies to the static + HTTP backends; the IMS/AKA paths use single-use HSS vectors. |
STIR/SHAKEN (Caller-ID Attestation)¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Sign — Authentication Service | Implemented | stir.sign(), stir: signing block |
ES256 PASSporT + RFC 8224 Identity header (RFC 8225, ATIS-1000074) |
| Verify — Verification Service | Implemented | stir.verify(), stir: verification block |
x5u fetch + full cert-chain validation to STI-CA anchors, sets verstat |
| Attestation levels A/B/C | Implemented | stir.sign(attestation=…) |
ATIS-1000074 §5.2.3; default via default_attestation |
Diverted-call PASSporT (div) |
Implemented | stir.sign_div() |
RFC 8946 — forwarded/retargeted calls |
| Cert chain + freshness | Implemented | stir.verification.freshness_secs, trust_anchors |
EC P-256 chain to STI-CA root; PASSporT iat window |
| Permissive rollout mode | Implemented | stir.verification.permissive |
x5u/infra failures → No-TN-Validation instead of …-Failed |
| x5u certificate cache | Implemented | stir.verification.cache_ttl_secs |
In-memory; honours Cache-Control: max-age |
verstat stamping |
Implemented | stir.apply_verstat() |
ATIS-1000074 §5.3.1 — P-Asserted-Identity / From |
| RCD (Rich Call Data) | Planned | Caller name/logo PASSporT — follow-up | |
| OCSP/CRL revocation, RSA STI-CA | Planned | EC P-256 chains only in v1 |
Security¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Rate limiting (per source IP) | Production | security.rate_limit |
PIKE-style fixed-window per-source-IP limiter. More than max_requests within window_secs → ban the source for ban_duration_secs (default 3600); every further request is dropped silently (no response — no fingerprinting). Enforced in the dispatcher on every inbound request before transaction/dialog/script processing, via the process-global crate::security::SecurityFilter (opt-in, installed only when configured). trusted_cidrs are exempt. 60s prune bounds the maps under scanner churn. Metric: siphon_rate_limited_total. Unit- + integration-tested (security::tests, tests/integration/security_tests.rs) |
| Scanner UA blocking | Production | security.scanner_block |
Drops any inbound request whose User-Agent matches a configured signature (case-insensitive substring — sipvicious, friendly-scanner, VaxSip, sipcli, …). Silent drop in the dispatcher (no response), trusted_cidrs exempt, same SecurityFilter path as rate limiting. When failed_auth_ban is also configured, a match over a connection-oriented transport (TCP/TLS/WS/WSS/SCTP — source validated by the handshake) escalates to a strong-weight auto-ban so the scanner's other probes are dropped at the ACL too; a match over UDP is only dropped (spoofable source → no reflected ban). Metric: siphon_scanner_blocked_total. Unit- + integration-tested |
| Trusted CIDRs (bypass rate limit + scanner block) | Production | security.trusted_cidrs |
Sources matching any CIDR bypass both the rate limiter and the scanner-UA block in SecurityFilter (own infra: AS/trunks/monitoring). Also exempted by failed_auth_ban's auto-ban store. Invalid CIDRs are ignored. |
| Trust gateways (provisioned carriers exempt from abuse controls) | Implemented | security.trust_gateways |
Off by default (behaviour unchanged). On: a source is trusted if trusted_cidrs holds it or a gateway group admits it, for every abuse control — rate_limit, scanner_block and its ban escalation, failed_auth_ban, connection_limits, APIBAN ingest. "Admits" has one definition, DispatcherGroup::admitted_sources (resolved destination addresses + source_networks), which from_gateway() and the kernel allow set already use, merged across siphon.yaml, script and gateway.backend groups into an immutable snapshot (gateway::view) of sorted, non-overlapping ranges: ArcSwap, lock-free, O(log n) binary search. Rebuilt off the request path wherever membership changes (group add/remove — so every reconcile and gateway.add_group() — each probe cycle's re-resolve, the allow-set floor tick). A publish that admits a new range lifts any auto-ban and evicts any APIBAN entry on it, userspace and kernel set alike, with a warn naming the group; a carrier removed from the source is back under every policy at that reconcile, counted from zero. A trusted gateway over rate_limit is allowed and logged once per window (its count is kept apart, so removal never inherits it; pruned with the rest). Unit-tested in gateway::view::tests, gateway::tests (static/script/source groups, agreement with from_gateway), security::trust::tests (each control on and off, with positive controls; leak test on the exempt windows), apiban::tests, and gateway::source::tests (eviction at the provisioning reconcile, back under policy at the removing one). |
| Failed auth ban (auto-ban) | Production | security.failed_auth_ban |
Per-source-IP auto-ban for toll-fraud scanners, fed by weighted failure signals so high-confidence abuse bans faster than a bare probe (strong_signal_weight, default 3, vs weight 1). Weight-1 (low-confidence): an auth challenge (401/407) not followed by a success; a non-ACK INVITE server-transaction timeout (RFC 3261 §17.2.1 Timer H); a failed/timed-out TLS handshake (kept weight-1 deliberately — a peer that distrusts the chain, offers no usable cipher, or is an L4 probe is benign, and weighting it as abuse turns a certificate rollover into a ban wave). Strong-weight (high-confidence): present-but-invalid digest credentials or a forged/stale/replayed nonce (kept weight-1 over UDP, where the source is spoofable → reflected-ban-safe); non-SIP bytes on a TCP/TLS stream, decided from the connection's first line before the SIP framer sees it (a complete HTTP probe such as a /phpinfo.php scan, an upgrade request on a listener that is not half of a tls+wss mux, a TLS record on the plaintext port, binary garbage, over-long header block — never an incomplete-but-plausible frame, a connect-and-close L4 health check, or a CRLF keepalive); a rejected WebSocket upgrade on a WS/WSS port, where the peer completed the transport handshake and then sent a non-upgrade request (RFC 7118 §5 gives a conforming client no way to produce one — so an external HTTP uptime monitor aimed at a WS/WSS port bans itself and belongs in trusted_cidrs); a scanner_block User-Agent hit over a connection-oriented transport (UDP scanner UAs are dropped, not banned). A successful auth resets the source's count, so a legit challenge→succeed (or stale-nonce retry) client never accumulates. threshold weighted failures within window_secs → ban for ban_duration_secs (default 10 / 600 / 3600). The ban expiry slides: a further signal from an already-banned source pushes it out to a full ban_duration_secs from that signal (previously such signals were discarded, so a scanner that kept hammering through its ban still walked out on the original schedule), capped by max_ban_duration_secs from the instant the ban was raised (default 24 × ban_duration_secs, clamped up to at least ban_duration_secs) — the cap is what keeps a source in a retry loop, which behind CGNAT speaks for a whole NAT, from being banned indefinitely. The rate_limit ban deliberately does not slide (capacity verdict, not intent). trusted_cidrs are exempt (own infra: BGCF/trunks/monitoring/health-check LBs) — client-transaction (relay-target) timeouts are also deliberately not counted, so a non-answering trunk is never banned. Enforced at accept/recv on every transport via TransportAcl::is_allowed (dropped before any SIP parsing), and re-checked once a TLS/WebSocket handshake completes so the sibling connections of a burst that were already past accept() when one of them tripped the ban die with it. Both of those are per connection, so a newly raised ban also closes every stream connection the source already holds open (TCP/TLS/WS/WSS/SCTP, one INFO line each, crate::transport::disconnect): a client that never reconnects used to keep the connection it had and go on presenting rejected credentials on it for the whole ban — observed in production as REGISTERs reaching the script for two minutes after the ban, each sliding the expiry further out, and a correct password accepted and stored eight minutes in with the source still banned — while every other client behind that address was refused at accept(). UDP never had the gap (the ACL runs per datagram). Regression-tested in transport::stream::tests::a_ban_closes_the_connection_its_source_already_holds_open (over one connection: three rejected credentials, the connection closed, the reconnect refused until expiry, and record_success not lifting a ban), transport::ws::tests::a_ban_closes_a_websocket_the_banned_ue_is_holding_open, and security::tests::every_ban_raising_path_closes_the_source_open_connections. Process-global store (crate::security::AutoBanStore, opt-in), lazy ban-expiry + 60s prune. Metrics: siphon_banned_ips, siphon_auth_failures_total, siphon_credential_failures_total, siphon_handshake_failures_total, siphon_malformed_messages_total. Unit-tested (security::tests, transport::tcp::tests, transport::stream::tests) plus end-to-end over a real loopback listener (an HTTP probe is closed and never dispatched; SIP in the same first segment still is) |
| APIBan integration | Production | security.apiban |
Community IP blocklist polling |
| IP ACLs (allow/deny CIDR lists) | Implemented | Transport-level ACL | |
| Kernel gateway allow set | Implemented | security.firewall.gateway_set |
Two nftables interval sets (gateways4/gateways6) holding every source the gateway groups admit — every resolved address of every destination (not just the one currently selected), every source_networks entry, and every security.trusted_cidrs entry, as written. The same view request.from_gateway() answers from, so the kernel and a script can never disagree about who is a gateway. Exists because a carrier that authenticates by source address has no registration and no outbound digest — the address list is the authentication — so a carrier provisioned at run time through gateway.backend was dialable while the kernel still dropped its answers: outbound appears to work while the inbound half is dead. siphon declares the sets and writes no rule for them in either manage_rule mode; the operator references them (ip saddr @gateways4 udp dport 5060 accept), because an accept inside siphon's own chain would also make a gateway immune to the ban drops, which is the operator's policy call. Published once during start-up before the listeners take traffic, on every reconcile and POST /admin/gateways/refresh, and on a 60 s floor tick that also re-resolves groups with probe.enabled: false — nothing else ever revisits their DNS, so a change of A record on one used to be invisible until a restart to routing and from_gateway() alike. A publish replaces the contents wholesale rather than diffing, so a carrier removed from the source stops being admitted; an unchanged view issues no netlink transaction at all. Survives the operator reloading the table that holds the sets, which is the table they have to name (security.firewall.table) to reference the sets at all, since an nf_tables set is scoped to its table: every wake-up reads back the kernel handle of the table and of each declared object, and on any difference re-runs the idempotent start-up declaration, drops the publish cache and republishes (warn + siphon_firewall_redeclared_total). Handles, not presence, because the operator's own file must redeclare the set its rules reference, so after a reload the set is back under its name and empty; a recreated table always gets a new handle from the per-namespace counter, while object handles restart inside it and can repeat. Also recovers from nft flush ruleset. Not seen: nft flush set (handle unchanged). Ban sets come back empty, so the re-declaration re-adds every live auto-ban and APIBAN entry in the same transaction, each with its remaining lifetime (permanent APIBAN entries stay permanent), skipping lapsed and now-trusted ones; the replay is split so each element list fits its 16-bit attribute, and the socket's send buffer is raised for a large batch. A replay the kernel refuses falls back to declaring the sets alone, with a warn (those bans stay userspace-only until they expire). Unit-tested on the batch contents (firewall::nftables::tests, firewall::tests). Covered by a second #[ignore] live-kernel test that reloads a table redeclaring every set in siphon's own order, so only the table handle can reveal it (verified to fail with the table handle left out of the comparison). Overlapping ranges are normalised away (CIDRs are disjoint or nested, so dropping the contained one is exact) because an interval set rejects an overlap. A CIDR rides as the element pair nft_set_rbtree takes — the element that opens the range plus an NFT_SET_ELEM_INTERVAL_END element keyed one address past its last — rather than being expanded to bare addresses or truncated to one. That is the encoding nft itself emits; the newer NFTA_SET_ELEM_KEY_END attribute reads like the obvious way to say the same thing and that backend rejects it with EINVAL, which is what the live-kernel test exists to catch and did. Unit-tested for the netlink encoding (host prefix, byte- and sub-byte-boundary v4 prefixes, v6, flush-carries-no-elements) and the range normalisation, plus the #[ignore] live-kernel roundtrip CI runs under unshare -rn, which publishes a mixed address/CIDR set and then replaces it. |
| Preloaded Route rejection | Production | Script logic | Anti-abuse for Route header |
NAT Traversal¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Symmetric response routing (rport, RFC 3581/6314) | Production | always on | Responses always go to the request source; no force_rport knob needed |
| Fix Contact (observed source) | Production | nat.fix_contact: true |
Rewrites the Contact on responses |
| REGISTER source capture | Production | automatic in registrar.save() |
Stored as Contact.received / Contact.flow for MT routing |
| Fix NATed Contact / REGISTER (script) | Production | request.fix_nated_contact() / fix_nated_register() |
Explicit REGISTER-side fixups |
| NAT keepalive (OPTIONS ping) | Implemented | nat.keepalive |
Configurable interval + failure threshold |
| CRLF keepalive (RFC 5626 §4.4.1) | Implemented | nat.crlf_keepalive |
TCP/TLS/pool connection keep-alive; outbound probe + inbound peer-ping/pong responder |
| Stale contact eviction on restart | Production | Core | Evicts connection-oriented contacts + on_change notify |
| Outbound flow tokens (RFC 5626) | Implemented | Via/Route flow tokens |
Media¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| RTPEngine integration (NG protocol) | Production | media.rtpengine |
Single or multi-instance; media.backend: rtpengine (default); rtpengine.answer(target, sdp=...) anchors a far side that is not a SIP agent (answer SDP from the script, rewritten SDP returned). received_from pins the party whose SDP a command carries, to that party's own source and by its own half of the profile: rtpengine.answer(reply) carries the source the reply arrived from, never the call= object's, also for a reply holding a delayed offer, and a re-offer from the callee or the caller's answer to one reads the callee's answer half and the caller's offer half. The command's shape is a separate question with the opposite answer: it is chosen for the party the engine's result is sent to, the callee of a dial under the offer half and its caller under the answer half, whoever offers, so a callee's hold on a profile whose halves differ in transport reaches each party in its own (MediaSession::party_shape; dispatcher::reoffer_ingress_tests a_callee_hold_is_relayed_to_each_party_in_its_own_transport for a re-INVITE and an UPDATE on an in-process native engine, delayed_offer_ingress_tests for the caller's answer to a delayed offer, script::api::rtpengine::answer a_command_is_shaped_for_the_party_it_is_sent_to_whoever_offers). SIPp-validated against a real rtpengine by the sipp-b2bua-refer CI job and run-tests.sh --b2bua (profile b2bua-rtpengine-reoffer): a caller offering RTP/SAVP with a crypto attribute and a callee on RTP/AVP, the callee holding and resuming with re-INVITEs, and each party asserting the transport, the crypto attribute and the direction of every SDP it is sent. The caller's 2xx to a callee's re-INVITE, answered through rtpengine.answer(reply), leaves the session's two tags naming the parties they named (the_callers_answer_to_a_callee_reoffer_leaves_the_sessions_tags_alone, rtpengine::session set_to_tag_never_records_the_offerers_own_tag) (script::api::rtpengine::answer a_reply_carries_its_own_source_and_never_the_callers, the_replying_party_is_pinned_by_the_answer_halfs_policy_also_on_a_delayed_offer, a_party_is_pinned_by_the_half_it_was_set_up_under_whichever_command_carries_its_sdp; through the Python methods against a stand-in native engine in a_reply_is_pinned_to_its_own_source_by_the_replying_partys_policy and a_reoffer_from_the_callee_is_pinned_by_the_callees_own_policy) |
| RTPEngine load balancing | Implemented | media.rtpengine.instances[] |
Weighted distribution |
| Native siphon-rtp backend (JSON/TCP) | Experimental | media.backend: siphon-rtp + media.siphon_rtp |
Persistent TCP control, auth handshake, reconnect, server-pushed DTMF/media-timeout events; same rtpengine API/profiles. The siphon-rtp engine is pre-release — use rtpengine in production. SIPREC/MPTY not yet supported on this engine. |
| Native siphon-rtp load balancing | Experimental | media.siphon_rtp.instances[] |
Weighted round-robin + per-call-id connection affinity; per-instance health probes (parity with rtpengine) |
| Classic rtpproxy backend (text/UDP) | Implemented | media.backend: rtpproxy + media.rtpproxy |
Classic U/L/D/V protocol with cookie correlation + idempotent retransmits; siphon-side SDP rewrite (multi-stream, held media); same rtpengine API/profiles. For migrating OpenSIPS/Kamailio/Sippy + rtpproxy. No prompts/DTMF/gating/SIPREC on this engine. Validate end-to-end against a live rtpproxy. |
| Classic rtpproxy load balancing | Implemented | media.rtpproxy.instances[] |
Weighted round-robin + per-call-id affinity; per-instance health probes (V) |
| Startup orphan media reap | Implemented | media.reap_orphans_at_startup (default false) |
Once at startup, before any listener binds, siphon enumerates the engine's live calls and deletes every one it has no session for — a restart loses siphon's session state while the engine keeps relaying, and nothing else removes those calls (the engine's media timeout only fires for a stream that went quiet). Matched on the engine-side call-id, not the SIP one, so a re-anchored live call is not mistaken for an orphan. Bounded by media.reap_limit (default 10000). Off by default: rtpengine's NG list is unscoped and answers with every call on the engine, so on a shared rtpengine an enabled reap deletes another node's live calls — enable only on a dedicated engine. The native siphon-rtp engine scopes list to the calling client and siphon claims a stable control identity from server.instance_id on every connection, so a restarted process is recognised as the same owner and a reap there can never reach another node's call even on a shared engine (an engine older than 0.8.0 ignores the claim: safe, but finds nothing). rtpproxy cannot enumerate: siphon warns at boot and skips. |
| Built-in profile: SRTP↔RTP | Implemented | srtp_to_rtp |
SRTP UE ↔ RTP core |
| Built-in profile: WS↔RTP | Implemented | ws_to_rtp |
WebSocket UE ↔ RTP core |
| Built-in profile: WSS↔RTP | Implemented | wss_to_rtp |
DTLS-SRTP/AVPF + ICE ↔ RTP |
| Built-in profile: RTP passthrough | Implemented | rtp_passthrough |
IMS-internal |
| Custom media profiles | Implemented | media.profiles |
User-defined NG flags |
| Media-plane IPv4↔IPv6 interworking | Implemented | media.profiles.<name>.{offer,answer}.address_family |
Pins the family (IP4/IP6) the engine allocates its relay endpoints in per direction, so a v6-only access leg can bridge to a v4 core; unset (default) = follow the offered SDP, unchanged single-family behaviour. rtpengine (dedicated address family NG key, not a flags token) + siphon-rtp (address_family control field); the classic rtpproxy backend cannot express it and is warned about at boot. Unknown values fail the config load. Wire-level tests on both backends (offer_carries_address_family_on_the_wire). Not yet validated against a live dual-stack engine. |
SDP manipulation (sdp namespace) |
Implemented | None | Parse/modify/apply SDP from Python scripts. Lossless on what it does not break out: a format that is a token rather than an RTP payload type (t38, webrtc-datachannel, *; RFC 8866 §5.14), its a=fmtp: line, an a=rtpmap: line it cannot read and a port/count all come back out as they went in, so a T.38, data channel, MSRP or floor-control section survives parse/apply and the paths that re-serialize SDP themselves (rtpproxy backend, control-plane re-bridge). An m= line with no protocol is no longer given RTP/AVP. Unit + integration tests, SDK-mirrored. |
| SDP attribute get/set/remove | Implemented | None | Session and media-level a= attributes |
| SDP codec filtering | Implemented | None | filter_codecs() / remove_codecs(), on RTP sections only: a udptl, UDP/DTLS/SCTP or TCP/MSRP section is left alone because its formats are not codecs. A stream left with no codec is rejected rather than emptied, port 0 with its first format kept (RFC 3264 §6, §8.2), never an m= line with no format. |
| SDP media section removal | Implemented | None | remove_media("video") |
| SDP attribute strip on B2BUA relay | Implemented | media.sdp_strip_attributes |
Removes the named a= attributes (session and media level, case-insensitive, with or without a value) from SDP relayed between B2BUA legs in both directions: the B-leg INVITE and its 401/407 and 422 retries, 18x, 2xx, relayed failures, re-INVITE and UPDATE and their answers, and siphon-originated re-INVITEs carrying the other leg's SDP (session refresh, transfer re-anchor, Replaces, controller bridge). Runs last, after the media engine rewrite; the engine still sees the peer's SDP. Only the application/sdp part of a multipart body. Names validated as RFC 8866 §9 tokens at config load; empty (default) skips the body entirely. Dispatcher-driven tests in dispatcher::sdp_strip_tests, including a fake NG engine for the ordering. Not yet SIPp-validated. |
Gateway Routing & Load Balancing¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Destination groups | Production | gateway.groups |
A destination's transport is udp (default), tcp or tls, case-insensitive, from the transport field or the URI's ;transport= parameter. Anything else is refused at config load naming the group, the destination, the token and which of the two keys carried it — and gateway.add_group() raises ValueError. A B-leg leaves over UDP or over a connection the outbound pool opens, and the pool opens TCP and TLS only, so sctp / ws / wss could never be dialled; effective_transport() also returned the explicit field verbatim, so transport: TLS matched neither "tcp" nor "tls" and carried a carrier's traffic, credentials included, in the clear over UDP. One parser (config::parse_outbound_transport) shared with the registrant entries, gateway.add_group() and the database/HTTP row check. |
| Round-robin algorithm | Production | algorithm: round_robin |
|
| Weighted algorithm | Implemented | algorithm: weighted |
|
| Hash-based algorithm | Implemented | algorithm: hash |
|
| SIP OPTIONS health probing | Production | gateway.groups[].probe |
Configurable interval + failure threshold |
| Priority-based failover tiers | Implemented | destinations[].priority |
|
| Dynamic group management | Implemented | Python gateway.add_group() / gateway.remove_group() |
|
| Gateway groups — database / HTTP source | Implemented | gateway.backend: database | http |
The destination groups are read from a source the controller owns and reconciled on refresh_secs. In place: an unchanged destination is carried over with its health, failure count and Retry-After cooldown intact, so a dead carrier stays dead across a poll instead of being resurrected every cycle, and a group with nothing different is not replaced at all (replacing one restarts its prober). An unreadable source keeps the current groups. Probe policy per group from the rows (probe, probe_interval_secs, probe_failure_threshold, probe_from_user, probe_from_domain; group-wide, first row carrying each wins, omitted = the 30 s default as before): a carrier that does not answer OPTIONS can be provisioned without being taken out of service by its own probe. A probe change replaces the group (the prober's period and From are fixed at spawn) but carries the destinations and their health; an unchanged one leaves the prober running. Switching probing off marks prober-downed destinations up, since with no prober nothing else would. probe_interval_secs: 0 is refused as a row (it would panic the prober task), and so is a transport — column or URI ;transport= parameter — siphon cannot dial: the column was already refused, but a URI-borne sctp / ws / typo used to be downgraded to UDP and put a carrier on the wire in the clear. The source owns only groups it created — gateway.groups and gateway.add_group() are never reconciled away. registers links a destination to an outbound registration: it inherits that credential, and require_registration gates selection on the registration being up. Versioned JSON contract typed in siphon_sdk.gateways, documented at docs/reference/gateway-api.md. Reconcile tests cover health preservation, group ownership and removal; not yet validated end-to-end against a live provisioning source. Read before the listeners take traffic, retried on a 250 ms backoff within a bounded start-up budget: the loop used to do its first fetch inside itself, so a node answered calls while its carriers were still being read, or — with the controller down — with none at all and one warn. The budget is bounded so a controller outage does not become an outage of every node that restarted during it; the node comes up loudly with zero carriers, which is a state to alert on, and POST /admin/gateways/refresh applies one the moment the source is back. A failed poll backs off from 250 ms capped at refresh_secs rather than sleeping the whole interval. siphon_gateway_source_last_success_timestamp_seconds and siphon_gateway_source_failures_total make it measurable (alert on the age of the first, the rate of the second; 0 = never read); both also appear on /admin/metrics.json under provisioning. A failure streak escalates warn → error from its second poll and is throttled to a line a minute — 720 identical warn lines is how a six-hour outage reads as normal. Unit-tested in gateway::source::tests (gives up inside its budget, backs off rather than spinning, counts every failure) and source_health::tests. |
| Destination up/down marking | Implemented | Python gateway.mark_up() / gateway.mark_down() |
|
| Source-membership predicate | Implemented | Python request.from_gateway() / call.from_gateway() |
ds_is_from_list() / ds_is_in_list() equivalent; IP-only match against all resolved group addresses, cached + refreshed on probe cycle. Trust signal on TCP/TLS/WS/WSS, direction hint on UDP |
| WhatsApp Business Calling (voice) | Implemented | examples/whatsapp_calling.{py,yaml} |
SIP-over-TLS trunk to wa.meta.vc (server-auth TLS, no mTLS), B2BUA both directions, outbound digest via call.set_credentials(), OPUS passthrough, SDES (built-in profiles) / DTLS-SRTP (whatsapp_dtls_*). Composed from existing features; live WABA validation pending |
Call Detail Records¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| CDR generation | Implemented | cdr: |
Auto-emit per call (cdr.auto_emit) — proxy + B2BUA, with duration + disconnect_initiator + Rf correlation; script cdr.write(request) (proxy) / cdr.write(call) (B2BUA) merges its extra fields into that record, and writes one of its own only when no auto-emit record is tracked. cdr.backends writes every record to several sinks (file / syslog / http) at once — one channel and writer task per sink, so a slow or failing sink cannot delay, block or drop another's records; a full channel drops for that sink alone, counted in siphon_cdr_dropped_total{sink}. No retry or durable queue for the HTTP sink: a file sink beside it is how a deployment gets durability. |
| REGISTER CDRs | Implemented | cdr.include_register |
With auto_emit, one CDR per registrar state change (reg_event) |
| File backend (JSON-lines) | Implemented | cdr.backend: file |
With rotation |
| Syslog backend | Implemented | cdr.backend: syslog |
UDP syslog |
| HTTP webhook backend | Implemented | cdr.backend: http |
POST with optional auth header |
| REGISTER event inclusion | Implemented | cdr.include_register |
Off by default |
| Script-injected extra fields | Implemented |
SIP Tracing¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| HEP v3 over UDP | Production | tracing.hep |
Homer integration. Every send is captured, the ones a background task makes included: the A-leg 2xx retransmitted until the ACK, the reliable 1xx retransmitted until the PRACK, and an IMS UE's protected REGISTER, as well as a request.relay(flow=…). Those used to bypass capture. A task resolves its capture once when it is armed (TaskCapture), and gets none without HEP, so a deployment without capture does no extra work per message. Dispatcher tests in dispatcher::retransmit_capture_tests; the IPsec UE REGISTER task needs installed SAs and is not unit-tested. |
| HEP over TCP | Implemented | tracing.hep.transport: tcp |
|
| HEP over TLS | Implemented | tracing.hep.transport: tls |
With CA cert + SNI |
| Custom agent ID | Production | tracing.hep.agent_id |
|
| Error log suppression | Production | tracing.hep.error_log_interval |
Configurable interval |
Metrics & Monitoring¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Prometheus endpoint | Production | metrics.prometheus |
|
| Request/response counters | Production | siphon_requests_total{method,direction} / siphon_responses_total{class,direction} |
Both directions, counted at the single inbound dispatch point and the single outbound send point. Counts wire events, so a retransmission counts each time — not a transaction count. Unknown methods bucket into OTHER (the token is attacker-controlled, so a series per token is a scrape-side cardinality DoS) |
| Active registrations gauge | Production | siphon_registrations_active |
|
| Active transactions gauge | Production | siphon_transactions_active |
|
| Active dialogs gauge | Production | siphon_dialogs_active |
Sum of siphon_proxy_dialog_sessions + siphon_b2bua_calls_active, both also exported separately — which side carries the load is what the roll-up cannot say |
| Active connections (by transport) | Production | siphon_connections_active{transport} |
Live inbound stream connections, released by an RAII guard so a cancelled task cannot leak the gauge upward. UDP is deliberately absent rather than zero — it is connectionless, so there is no connection to count |
| Script error counter | Production | siphon_script_errors_total |
Incremented alongside every Python handler-error log, through one helper, so counter and logs cannot disagree |
| Uptime gauge | Production | siphon_uptime_seconds |
Published by the dispatcher sweep, so a Prometheus-only deployment sees it without opening the dashboard |
| Admin API — health | Implemented | GET /admin/health |
Liveness/readiness probe. admin.listen must be an IP:port socket address (refused at load otherwise), and a listener that cannot bind ends the process instead of leaving it healthy with no admin API. Unit-tested in config::tests::admin_listen_* |
| Admin API — bans | Implemented | GET/DELETE /admin/bans |
List / lift auto-bans. Both 404 when security.failed_auth_ban is off, and security.banned_ips in /admin/metrics.json is null then, so an empty list only ever means "watching, nothing banned". Unit-tested (admin::tests::bans_*) |
| Admin API — retained logs | Implemented | GET /admin/logs, admin.log_tail.retain_level |
WARN+ always retained (warn_capacity); opt-in second ring down to retain_level (retain_capacity) so a call can be read back after it ended. Filters level / contains / call_id, paging limit + before=<seq>, retained_level in the answer. Bounded rings, leak-gated in log_tail::tests::steady_state_does_not_grow_the_retained_rings |
| Admin API — stats | Implemented | GET /admin/stats |
Aggregate counters |
| Admin API — registrations | Implemented | GET/DELETE /admin/registrations |
List, detail, force-unregister |
| Admin API — metrics snapshot | Implemented | GET /admin/metrics.json |
Curated JSON of live gauges/counters (SIP, traffic by method/class, memory, pyexec, Diameter per-command, per-instance media, control plane, charging sessions, SBI, IPsec, security) for the dashboard + custom tooling; browser diffs cumulative counters for rates. A subsystem this node has not configured reports null, never a zero — so "absent" and "idle" are distinguishable. rtpengine reports only where a media engine is configured (a media: block that names a backend, gives an engine block, or sets profiles/events); a block with only sdp_name/sdp_strip_attributes reports null like no block at all. An engine configured but unreachable still reports, with its instances down |
| Admin API — bearer auth | Implemented | admin.auth.token / admin.auth.protect_reads |
Constant-time bearer check; gates DELETE (and, with protect_reads, reads + /metrics). Unset = open (unchanged). A failed token feeds the auto-ban with the source address, over plaintext and TLS alike. protect_reads: true no longer costs the dashboard: the UI shell is served unauthenticated and every byte of data stays behind the token. The SPA fallback sits below the auth layer, so its own assets used to 401 and the dashboard could never bootstrap far enough to ask for one — which forced protect_reads: false on any node that wanted a UI, leaving the registration list (number, contact address, expiry) and the live call list readable to whatever the admin ACL admitted with no token at all. The exemption is a GET/HEAD under neither /admin nor /metrics, which is exactly the set that falls through to the SPA handler, so it cannot reach a data route; /admin/logs, /admin/capture and /admin/search stay always-gated. Unit-tested per route, both directions, and with the UI disabled. |
| Admin API — TLS | Implemented | admin.tls |
The same TlsServerConfig the SIP listeners take, including the rename-watching hot reload — a renewed certificate is served by the next handshake, so cert-manager and certbot need no restart — and verify_client/client_ca for mutual TLS, which is the stronger answer for a controller than a bearer token alone. Without it the token crossed the wire in the clear on every call and so did everything the API returned, which is why driving POST /admin/gateways/refresh from a controller meant an ssh tunnel. An unreadable certificate or key fails the config load naming admin.tls, rather than binding a listener that looks healthy and fails every handshake. Shares one TlsListener with the control plane (transport::tls_listener): the handshake runs on a task per connection, never inline in accept(), so one peer that connects and stalls cannot hold up the listener. Unit-tested end-to-end over a real socket — a GET /admin/health completed through rustls against a generated chain — plus config tests for the refusal. |
| Admin API — gateways | Implemented | GET /admin/gateways |
Per-group dispatcher status: each destination's health/weight/priority/address/transport/attrs plus its consecutive missed health-checks (checks_missed) and the group's failure_threshold, read from the shared dispatcher (no new state or probing). A group with an inbound_limit reports it with calls_active; one without reports null |
| Graceful shutdown teardown | Implemented | server.teardown_secs |
At the drain deadline, the calls still up are ended instead of being exited on top of: a BYE on both legs carrying Reason: Q.850;cause=16 (normal clearing — an orderly hangup the network chose, not a fault), the Ro CCR-TERMINATION, the Rf ACR-STOP, the media release, SIPREC stop and the CDR, through the same funnel b2bua.terminate uses. An unanswered call has every pending B-leg CANCELled (RFC 3261 §9.1 — removing a call only stops its leg actors and emits nothing on the wire) and its caller gets the 503 the drain already answers a new INVITE with. A call lasts minutes and drain_secs is seconds, so the deadline is the normal path on any restart taken with traffic up, not the exceptional one: before this it logged one line and exited, leaving the far side holding a channel until someone hung it up, anchors held to media timeout and charging sessions the OCS had authorised never closed. The pass then waits up to teardown_secs (default 5) for the work to land — the charging stops and the media delete are spawned and awaited nowhere, so exiting straight after issuing them would kill the Diameter round trips mid-flight. It deliberately does not re-route: the @b2bua.on_failure path a ring timeout runs would have an exiting node start dialling carriers. teardown_secs: 0 restores the previous behaviour exactly. Two stated limits: a BYE for a dialog whose peer has not ACKed siphon's 2xx is held (RFC 3261 §15) and goes out only if the ACK arrives, and proxy-mode calls cannot be torn down at all — ProxySession is transaction state, not dialog state (no callee To-tag, remote target, route set or CSeq), so a pure proxy node reports no active calls and drains instantly while proxied calls are up. Operator contract: the container runtime's stop timeout must exceed drain_secs + teardown_secs or its SIGKILL wins (Docker defaults to 10 s). Unit-tested against the real dispatcher (answered → two BYEs with the Reason; ringing → CANCEL + 503; a call another teardown owns → untouched; empty store → nothing on the wire), plus a SIPp acceptance test against the real binary and a real SIGTERM (scripts/run-tests.sh --shutdown), whose second arm proves teardown_secs: 0 still leaves both legs alone. |
| Admin API — gateway control | Implemented | POST /admin/gateways/{group}/{dest}/{up\|down} |
Manual mark-up/down of a destination (drain a bad carrier, then restore it); mutating, so it sits behind the bearer gate |
| Admin API — provisioning refresh | Implemented | POST /admin/registrants/refresh | POST /admin/gateways/refresh |
Re-read a backend: database / http source and reconcile now, so a controller applies a change on save instead of waiting out refresh_secs. Mutating, so behind the bearer gate. Answers the pass counts (added/updated/removed/rejected); 501 on a statically-configured node (nothing to re-read, not a failure) and 502 with the live set unchanged when the source cannot be read |
| Admin API — calls | Implemented | GET /admin/calls |
Active B2BUA calls, read from the dispatcher-owned call store: Call-ID, state, ring/talk duration, caller and dialed callee, per-branch status including the failure code and which branch won, per-leg transport and remote address, session-timer state, transfer state, controlling app, recording flag, and per-carrier LCR attempts (with dialed, so a local gateway/DNS fault is not misread as a carrier fault). Re-INVITE/UPDATE tracking pseudo-legs excluded. Empty on a proxy-only node |
| Web dashboard (embedded) | Experimental | ui cargo feature + admin.ui.enabled |
Operator UI baked into the binary, served same-origin on the admin listener: Overview / Calls / Registrations / Gateways / Signalling / Media / Control / Security / System. A metric whose subsystem is not configured renders as "not configured", never as a zero. Chart history survives a reload; only the visible view polls, and polling backs off while the tab is hidden. Compiled into the release Docker image by default; off for the plain cargo build, library consumers unaffected. Serving it logs an EXPERIMENTAL warning; feature-off + enabled: true warns and serves nothing |
Logging¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| JSON structured logging | Production | log.format: json |
|
| Pretty (human-readable) logging | Implemented | log.format: pretty |
|
| File logging | Production | log.file |
With logrotate support |
| Log level control | Production | log.level |
debug/info/warn/error |
| Non-SIP payload drop + parse-error suppression | Production | none | Keeps a well-behaved peer's non-SIP traffic from burying the parse errors that matter. Three payload shapes are dropped at TRACE before the parser, counted on siphon_non_sip_datagrams_dropped_total{reason} and never scored toward the auto-ban: whitespace-only (RFC 3261 §7.5 / RFC 5626 §4.4.1, any transport), an all-NUL UDP datagram (a vendor NAT keepalive no RFC defines — RFC 5626 §4.4.1 is CRLF or STUN — sent at a registered contact every few seconds for the life of the registration), and a UDP datagram under the 14-byte grammar floor for a SIP start line (the shorter of Status-Line "SIP/2.0" SP 3DIGIT SP CRLF = 14 and Request-Line with a one-character method and a three-character absoluteURI = 15, so it is the point below which no start line can complete). The bound is the grammar floor, not a guess about the peer, so a malformed-but-plausible message still reaches the parser and still warns. Everything the parser does reject is warned about once per source per 60 s, with the next line past the window carrying the count that went unlogged; the periodic sweep flushes an outstanding count when a source goes quiet and forgets it, and the table is capped at 4,096 sources (past the cap every error logs rather than the table growing). Unit-tested (dispatcher::inbound_filter — grammar floor, each drop shape, per-transport asymmetry, rate-limit windowing, source-ceiling, steady-state table drain; dispatcher::inbound_drop_tests — the same through handle_inbound, asserting the counter and that the parser was or was not reached) |
Python Scripting¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Script loading | Production | script.path |
|
| Hot-reload via inotify | Production | script.reload: auto |
Scoped to the code this process runs. It used to reload on any .py file in the script's directory or an include_paths entry, so two siphon processes sharing a directory reloaded each other — and since a reload re-executes the script in a fresh namespace and purges its helpers from sys.modules, the process that had not changed lost its module-level state (measured: a routing-script deploy emptied a Diameter charging bridge's per-session map, and every call answered afterwards lost the answer instant its records are built from). siphon now reloads for the script's own file and for helper modules the running script has imported, read from sys.modules at event time rather than snapshotted at compile time so a lazily-imported helper still triggers one; a module nothing has imported is not a trigger and needs none, since there is no stale copy to replace. Events within 250 ms are coalesced into one reload — the previous 50 ms sleep delayed each event and then reloaded once per event, so an editor's save or a deploy writing three helpers cost a recompile each. Watching is non-recursive, so a helper laid out as a package (lib/mypkg/__init__.py) does not trigger a reload — unchanged, and now documented rather than implied. Unit-tested in script::engine::tests (own file, imported helper, un-imported sibling, lazily-imported helper) and script::watcher::tests (event kinds, coalescing); end-to-end in scripts/run-tests.sh --reload. Every load gets a siphon module of its own, installed whole: handlers still running from the replaced script keep the module they started with, the reloaded script gets the new one, and a failed reload puts the old one back (it used to re-run the package inside the live module, rebinding namespaces to stubs under running handlers). A reload takes over the custom metrics the replaced script declared, values kept, when type, labels and buckets are unchanged (it used to fail with "already registered" for any script declaring metrics). Tested in script::engine::tests::a_reload_leaves_a_running_handler_the_siphon_module_it_started_with, a_failed_reload_keeps_the_siphon_module_of_the_script_it_keeps, a_reload_redeclares_the_metrics_its_script_declares. |
| Hot-reload via SIGHUP | Implemented | script.reload: sighup, and kill -HUP under reload: auto too |
Was Implemented while nothing implemented it: ReloadMode::Sighup was documented as "only reload on SIGHUP", its sole effect was to switch the inotify watcher off, and no SignalKind::hangup() handler existed anywhere in the tree — so the mode meant never reload, silently, and kill -HUP terminated the process on the default disposition. The worst of the three possible behaviours, since an operator picks the mode precisely to control when module state is wiped. SIGHUP now reloads the script, installed in both modes (POST /admin/script/reload already worked regardless of mode, so a signal that worked in one and not the other would be the same trap). Under sighup the watcher stays off, so the signal remains the only file-driven trigger. Gated by scripts/run-tests.sh --reload against the real binary — signal delivery is not reachable from a unit test — which also asserts that a file write under sighup still does not reload. Not a per-PR CI job, the same standing as the --nohandler fallback regression it is modelled on. |
| Proxy handlers (on_request/on_reply/on_failure) | Production | @proxy.* |
on_request + on_reply proven. @proxy.on_reply takes the same optional method filter as on_request ("INVITE|UPDATE"), matched against the relayed request's method; @proxy.on_register_reply is shorthand for on_reply("REGISTER") (it previously registered but was never dispatched). Dispatch-tested through run_reply_handlers in dispatcher::proxy_reply_filter_tests, SDK-mirrored in sdk/tests/test_reply_filter.py. Per-relay request.relay(on_reply=…) / request.relay(on_failure=…) callbacks now fire correctly on the free-threaded build: the dispatcher response path lifted the stored on_reply/on_failure Py<…> callbacks out of the session read-guard with a bare Clone on a Python-executor worker that was not inside a Python::attach scope. Under free-threaded CPython (3.14t, pyo3 0.28) Py::clone panics ("Cannot clone pointer into Python heap without the thread being attached") unless the thread is attached, unwinding the worker mid-relay — which truncated the in-flight relayed request and failed every call that armed a per-relay callback (blocked reply-driven MMTel behaviours: CFNR / busy-on-200 OK marking). Fixed by ProxySession::clone_relay_callbacks, which clones through a Python token (clone_ref) under Python::attach, matching the request path's discipline (script::handle::call_handler). Regression-tested in proxy::session::tests::clone_relay_callbacks_from_unattached_worker_thread (clones the callbacks from a freshly spawned, never-attached OS thread). |
| B2BUA handlers | Production | @b2bua.* |
on_invite, on_early_media, on_answer, on_failure, on_bye, on_refer. What @b2bua.on_failure decides is carried out on every path a call fails on (branches exhausted, ring timeout, INVITE never sent, no LCR carrier routable, an answer on_answer refused, a caller Require the call cannot honour, answered 420 Bad Extension with Unsupported per RFC 3261 §8.2.2.3 before any B-leg goes out, dispatcher::bad_extension_tests; the same 420 ends the call without on_failure where no script routing decision applies: siphon answering the call itself, an answer-first handover, the control plane's dial/route, and a Replaces takeover, dispatcher::uas_bad_extension_tests): call.reject() sets the caller's response, call.dial()/fork()/route() route the call again (at most 10 times per call), call.handover() hands it to a control app, call.answer() keeps it. Previously the decision was run and ignored, and the call ended with its failure. on_failure and on_cancel run once per call outcome even when two messages about it are handled at once (a retransmitted final response or CANCEL, the ring timeout firing while the last branch fails), dispatcher::b2bua_conclude_once_tests. Unit-tested in dispatcher::b2bua::on_failure::tests (classification) and b2bua::actor::tests (re-route and answer rewind); SIPp-validated by the sipp-b2bua-on-failure CI job (re-dial after a 486, a 480 replacing a 486, re-dial after the ring timeout, re-dial of an INVITE that never left, re-dial after a refused answer, an answer from the handler), and the handover by the sipp-control job (a dial that never resolves, handed by the failure handler to a control app that answers it). |
| Registrar hooks | Production | @registrar.on_change |
|
| Auth API | Production | auth.require_digest() etc. |
|
| Gateway API | Production | gateway.select() etc. |
|
| Cache API | Production | cache.fetch() |
Redis-backed |
| Cache list / TTL / existence ops | Implemented | cache.list_push/list_pop_all/list_len/list_len_sum/expire/exists |
Redis-backed FIFO queue ops (atomic LRANGE+DEL drain), single-list length (LLEN), prefix-summed depth (SCAN+pipelined LLEN, TTL-expiry-truthful), per-key TTL, presence check; degrades silently when Redis is unreachable |
| Presence API | Production | presence.* |
Used for reg-event SUBSCRIBE/NOTIFY |
| Outbound SUBSCRIBE (RFC 6665 watcher) | Implemented | proxy.subscribe_state.send/find/refresh |
Originate SUBSCRIBE, capture dialog state from 200 OK, correlate inbound NOTIFY by tags |
| SUBSCRIBE dialog handle | Implemented | SubscribeHandle properties, reload() |
Properties read this instance's dialog state — live, never a snapshot, and never a network wait, so an async def handler cannot pin its asyncio driver on one; a reaped dialog raises LookupError and await handle.reload() is the explicit L2-cache re-read for the cross-replica case |
| SUBSCRIBE notifier accept | Implemented | proxy.subscribe_state.accept(), handle.notify() |
The 200 accept() stages leaves before any NOTIFY or terminating NOTIFY the same handler awaits (RFC 6665 §4.1.2.3): the handle holds them until the handlers return, whatever thread the coroutine ran on. Later sends go out immediately. Regression-tested through handle_request with a real async def script (the_initial_notify_leaves_after_the_200_that_accepted_the_subscription) |
| Reginfo XML parser (RFC 3680) | Implemented | presence.parse_reginfo(xml) |
Watcher-side parser for application/reginfo+xml NOTIFY bodies |
| Lawful intercept API | Implemented | li.* |
|
| Logging API | Production | log.* |
|
| Async handler support | Production | Auto-detected by runtime | |
| Custom metrics API | Production | metrics.counter/gauge/histogram |
Script-defined Prometheus metrics |
| Timer routes | Implemented | @timer.every(), timer.set()/cancel() |
Periodic callbacks via Tokio; one-shot cancellable timers keyed by string |
| Mock SDK for testing | Implemented | siphon-sip (imports as siphon_sdk) |
Test scripts without Rust binary. The digest checks (auth.verify_digest / require_www_digest / require_proxy_digest / require_digest) do the real RFC 7616 §3.4 arithmetic when the script supplies password= or ha1= — H(A1), H(A2) and the response under the algorithm the Authorization header names (MD5 / SHA-256 / SHA-512-256, with and without qop=auth), with ha1 bound to its own hash per §3.4.3 — so a wrong credential returns False in the mock exactly as it does on a node. They previously answered from the mock's preset allow flag, which silently cost every script delegating verification to the engine its accept/reject assertions. With neither keyword given the flag still governs; there is no credential source in a mock. Nonce replay is not checked by the mock (the engine checks it first) — that is validate_nonce's own test. sdk/tests/test_auth_verify_digest.py, whose vectors are built with hashlib in the test so the test and the mock are not the same arithmetic. |
| Extension API (host namespaces, tasks, module extensions, custom handler kinds) | Implemented | extensions:, register_namespace/register_namespace_with/register_module_extension/register_task, _siphon_registry.register("custom.kind", …) |
Open extension surface for custom transports / sinks; ScriptHandle::handlers_for + call_handler dispatch into script handlers from host extensions. register_module_extension mounts a multi-namespace surface (several namespaces + shared types + an exception) onto the siphon module in one hook — what the SIGTRAN module uses for ss7/gsm_map/gsm_cap/inap. Shipped modules: smpp, http, sigtran (see extensions.md) |
| Elastic handler pool (grow + bounded queue + watchdog) | Production | script.sync_pool_size / sync_pool_max, executor_queue_capacity, handler_stall_abort_secs |
Pool grows core→max under blocking load and never reaps (no wedge, no heap leak); bounded queue load-sheds at the cap; deadlock-aware liveness watchdog aborts (→ supervisor restart) on zero forward progress while work is pending, at any pool fill (catches low-concurrency deadlocks). Blocking Rust-API calls release the interpreter (py.detach) to avoid the free-threaded GC stop-the-world deadlock. Regression-guarded by pool_grows_under_blocking_load, detached_blocking_does_not_stall_gc, and run-tests.sh --http-auth. Metrics: siphon_pyexec_*. See handler-execution-model.md |
Dialog Management¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Memory backend | Production | Default | In-process, ephemeral |
| Redis backend | Implemented | dialog.backend: redis |
Persistent across restarts |
| PostgreSQL backend | Implemented | dialog.backend: postgres |
Named Cache¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Redis-backed cache | Production | cache[].url |
|
| Local LRU tier | Implemented | cache[].local_ttl_secs |
Two-tier: local + Redis |
Presence¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| SUBSCRIBE/NOTIFY (RFC 6665) | Production | Python presence API |
reg-event package; presence.terminate() + auto-GC on terminated NOTIFY drops dialog state per RFC 6665 §4.4.1 |
| PIDF (RFC 3863) | Implemented | ||
| Resource List Server (RFC 4662) | Implemented | ||
| Watcher Info (RFC 3857/3858) | Implemented |
Server Identity¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Custom Server header | Production | server.server_header |
|
| Custom User-Agent header | Production | server.user_agent_header |
Transaction Timers¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Non-INVITE timeout | Production | transaction.timeout_secs |
|
| INVITE timeout | Production | transaction.invite_timeout_secs |
How long proxy state for an INVITE is kept once no branch of it is still owed a final response. A session with a pending branch is never swept, however long the call rings (dispatcher::proxy_timer_c_tests::a_call_answered_after_ringing_past_the_transaction_timeout_still_connects). |
| Timer C (RFC 3261 §16.6 step 11) | Implemented | transaction.timer_c_secs (default 181) |
A proxied INVITE that has a provisional and no final response is CANCELled when Timer C runs out (§16.8), counted from the INVITE and again from each 101-199 (§16.7 step 2); its 487 is the branch's final response, and a branch silent for 64T1 after the CANCEL is given up as a 408. A provisional resets it by storing its arrival time, so the timer wheel is only touched once the INVITE has outlived Timer B. Unit-tested in transaction::timer_c_tests (the default and the configured value reaching the timer, the hand-over from Timer B, reset by 101-199 and not by 100, the CANCEL, the 64T1 bound, no second CANCEL, a final response stopping it) and through the dispatcher in dispatcher::proxy_timer_c_tests (the CANCEL and forwarded 487, the reset, the 408, a fork branch counted by the aggregator, and nothing left behind). |
DNS Resolution¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| SRV lookup (RFC 3263) | Implemented | Core | With A/AAAA fallback; weighted-random RFC 2782 selection per call |
| A/AAAA load distribution (RFC 3263 §4.2) | Implemented | Core | Fisher-Yates shuffle on every A-only resolution so callers picking .next() distribute uniformly across equal-cost records |
| NAPTR support | Implemented | Core | |
| ENUM (RFC 6116) | Implemented | Core |
3GPP / IMS / Telco¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| Diameter Cx (HSS auth) | Production | auth.backend: diameter_cx |
MAR/SAA, SAR/SAA, UAR/UAA, LIR/LIA |
| Diameter Sh (HSS user data) | Production | diameter |
sh_udr for repository data; inbound PNR (profile push) handled via @diameter.on_request (req.command_name == "PNR") |
| Diameter Ro (online charging) | Implemented | diameter, ro: |
CCR/CCA (RFC 8506 / TS 32.299) with the correct RFC 8506 AVP codes on the wire. Two flows: SCUR for voice (CCR-INITIAL reserve at setup → CCR-UPDATE re-auth on the OCS-granted quota → CCR-TERMINATION on BYE, with mid-call disconnect on 4012/Final-Unit-Indication, 4011 free-tier, and fail-open/closed CCFH) and IEC for SMS/RCS (one-shot CCR-EVENT, DIRECT_DEBITING). Mandatory Service-Context-Id, Multiple-Services-Indicator, single-MSCC-vs-command-level per §5.1.2, parse_cca reads grants at both levels. B2BUA-only enforcement (ro: config auto-emits on the B2BUA path — reserve at answer, re-authorize, disconnect on denial, terminate on BYE): cutting a live call needs to own the session, which is why 3GPP triggers Ro at the AS/MMTel-AS, not the P-CSCF (no proxy-mode auto-emit; the raw diameter.ro_ccr_* scripting API works in any mode for manual use / one-shot IEC). siphon_ro_sessions leak gauge, plus siphon_ro_denials_total{result_code} (a call refused credit at setup) and siphon_ro_credit_teardowns_total{reason} (one cut off mid-call) — neither moved any counter before, since the CCR/CCA round trip succeeds and no SIP error fires; the reason="no_teardown_hook" series is the alertable one, meaning credit ran out with nothing wired to enforce it. Carrier attribution: Outgoing-Trunk-Group-Id (TS 32.299 §7.2.71) is stamped on the session at every LCR attempt that reaches the wire and again at the answer, so it means "the carrier that answered, or the last one dialled if none did" and reaches a mid-call re-authorization and the final record alike — absent only when no carrier was ever dialled. It used to be stamped at the 2xx alone, so every unanswered call (a caller hanging up during ringing above all, then busy, ring timeout, a carrier's own 5xx after failover) produced a CCR-TERMINATION naming nobody, and a per-carrier answer-seizure ratio computed from the charging feed read 100 % for every carrier. The CDR feed had the same hole on the same calls — the caller-cancel path wrote its 487 without the route's cdr_fields or lcr_attempts — so there was nothing to reconcile against; both are stamped there now. Validated against an in-process mock OCS (SCUR reserve→re-auth→deny→teardown→drain, IEC grant/deny, denial + teardown counters, and the carrier on the final record read off the OCS side for a cancelled call, a failover cancelled on its second carrier, an answered failover and a sequence that reached no carrier — dispatcher::carrier_attribution_tests); a CGRateS docker charging profile + SIPp scenarios author the live acceptance test. |
| Diameter Rf (offline charging) | Implemented | diameter, rf: |
ACR/ACA wired through diameter.rf_acr_start/interim/stop/event (TS 32.299 §6.2.2) — kwargs-style Python API, mandatory AVPs (Service-Context-Id, Event-Timestamp, User-Name, Subscription-Id (0..n, subscription_id / subscription_id_type accept one value or a list), Termination-Cause, Acct-Interim-Interval), full IMS-Information sub-AVPs (User-Session-Id, Time-Stamps, Inter-Operator-Identifier, Application-Server, IMS-Visited-Network-Identifier), TS 32.260 IMS Service-Context-Id default. SMS-Information envelope (TS 32.299 §7.2.79) — passing any SMS-specific kwarg (originator_address, recipient_address, sm_message_type, sms_node, sm_user_data_header, reply_path_requested, sm_service_type, sms_result, SCCP/Client/MTC-IWF Address fields, sm_discharge_time, data_coding_scheme, …) switches the wire to Service-Information → SMS-Information so CDR collectors render calling/called party + message type on the SMS tab; can coexist with IMS-Information for hybrid records. rf: config block + RfChargingService runtime emits ACR-EVENT automatically on registrar state change. CDR auto-stamps rf_session_id / rf_result_code from auto-emitted records. B2BUA + proxy ACR-START/INTERIM/STOP auto-emit on call lifecycle: START is emitted on the 200 OK (TS 32.260 §5.2.2) with Time-Stamps carrying the INVITE and the final response separately so a CDF can derive ring time, an ACR-EVENT with a negative Cause-Code reports unsuccessful session establishment (§5.2.2.1), and Calling-Party-Address + Subscription-Id repeat once per asserted identity for a multi-valued P-Asserted-Identity. |
| Diameter Rx (policy/QoS) | Production | diameter |
AAR/AAA, STR/STA; inbound RAR/ASR handled via @diameter.on_request (req.command_name). diameter.rx_aar(media_components=[…]) takes a list of TS 29.214 §5.3.7 MediaComponent dicts with per-flow IPFilterRules + Flow-Usage (RTCP marker) — pair with qos.media_flows_from_sdp(offer, answer, direction) to derive the full 5-tuple from an SDP offer/answer rather than emitting a wildcard permit in 17 from <UE> to any that any non-permissive PCEF would either drop or open globally. diameter.rx_aar(specific_actions=[…]) subscribes to PCRF event reports (TS 29.214 §5.3.13), one Specific-Action AVP per int, e.g. 9 INDICATION_OF_FAILED_RESOURCES_ALLOCATION; a value TS 29.214 does not define (including the void 0 and 5) raises ValueError, in siphon and in the SDK mock. rx_aar and the crate's rx::send_aar share one encoder (rx::encode_aar), so every AAR carries Rx-Request-Type (INITIAL_REQUEST, or UPDATE_REQUEST on a reused session_id; V-bit only, per table 5.3.1) and the configured Destination-Host. SpecificAction values corrected to TS 29.214 (IP-CAN_CHANGE is 6, not 7; 7/8/9 and 10-21 added). The full AAR, including the Specific-Action values, is checked against Wireshark's Diameter dissector (scripts/validate_rx_aar.sh) and pinned by known-answer tests; the subscriptions are not yet validated against a live PCRF. |
| Diameter S6c (SMS-over-Diameter, SMSC↔HSS) | Implemented | diameter |
s6c_srr to discover served-node, s6c_rsr for delivery status; inbound ALR (HSS reachability alert) handled via @diameter.on_request (TS 29.336). MSISDN / SC-Address / SGSN-Number / MME-Number-for-MT-SMS encoded as ISDN-AddressString (TS 29.002 §17.7.8 — ToN/NPI 0x91 + TBCD digits); inbound parser is lenient on missing ToN/NPI prefix for non-conformant peers. |
| Diameter SGd (SMS-over-NAS, SMSC↔MME) | Implemented | diameter |
sgd_tfr to deliver SMS-DELIVER TPDU to UE; inbound OFR (MO-SMS) handled via @diameter.on_request (TS 29.338). SC-Address on the wire uses ISDN-AddressString (TS 29.002 §17.7.8), matching S6c. |
| Diameter S6a (MME↔HSS, LTE attach/auth) | Implemented | diameter.s6a_air/s6a_ulr/s6a_purge_ue (client); @diameter.on_request + req.answer() (server) |
TS 29.272 — client: AIR/AIA (E-UTRAN vectors RAND/XRES/AUTN/KASME, SQN resync), ULR/ULA, PUR/PUA. Server (HSS role): siphon transports inbound AIR/ULR/PUR to @diameter.on_request; the script builds the answer with req.answer(code) + grouped-AVP construction. siphon does NOT implement S6a semantics or Milenage — the script owns subscriber data + auth-vector crypto (see examples/hss_s6a.py). Relayable by a server-mode script. Dictionary AVPs 1400–1450/1635 + command codes 316–324. |
| Diameter generic answer + grouped AVPs (server) | Implemented | req.answer(result_code), DiameterRequest/DiameterAnswer.{get,set,insert}_avp |
Application-agnostic inbound serving: build a local answer envelope and construct/read arbitrarily nested Grouped AVPs from Python (list of (code, value[, vendor]) child tuples; values may nest). Lets a script serve any Diameter application (HSS/PCRF/OCS) on the inbound listener — siphon transports, Python decides. |
| Diameter serve-on-outbound (dial-out + serve) | Implemented | diameter.connect_to |
A server NF that initiates the connection (e.g. an HSS dialling an upstream) but answers the requests relayed back over it. siphon sends the CER, then routes inbound requests to @diameter.on_request exactly like the listener path — transport direction is independent of request direction (RFC 6733 §2.1). Works without diameter.listen. TCP + SCTP. |
| Diameter generic API (spec-name addressing) | Implemented | diameter.send_request("Send-Routing-Info-for-SM-Request", application="S6c", **avps) (originate); @diameter.on_request + req.command_name (serve) |
Outbound origination by spec name (AVPs encoded by dictionary type, snake_case ↔ kebab-case kwargs, 3-letter acronym aliases SRR/ALR/TFR/…). Inbound serving is the single unified @on_request hook — the old per-command @on_command was removed. |
| Diameter peer management | Production | diameter.peers |
Failover + round-robin across HSS/PCRF peers. Observability: siphon_diameter_peer_up{peer} reports each configured peer 0/1 by name (every peer is published at 0 before its first connect attempt, so one that has never come up reads as down rather than being absent), siphon_diameter_answers_total{command,result_code} counts what peers actually answered — siphon_diameter_request_errors_total covers transport failures only, so a peer that is reachable and refusing reads as zero errors there — and siphon_diameter_request_duration_seconds surfaces per-command round-trip mean/p95 on the admin snapshot. |
| Diameter server mode | Implemented | diameter.listen, diameter.clients, diameter.servers, @diameter.on_inbound_cer, @diameter.on_request, @diameter.on_reply, @diameter.on_request_completed |
Accepts inbound Diameter (TCP + SCTP), runs CER/CEA + the DWR/DWA watchdog, and dispatches each inbound request to Python — siphon transports, the script decides (answer locally or relay). Two Rust-only admission gates (source-IP CIDR ACL + Origin-Host validation, both before any Python), lossless AVP tree (DiameterRequest/DiameterAnswer get/set/remove/insert/iter), req.forward_to(peer) relay with Route-Record loop detection (3005) + per-call timeout, @diameter.on_reply for central answer-AVP rewrite (topology hiding, Origin/Result-Code mapping), diameter.peer_pool(target) (round-robin / weighted / sticky over state-as-truth liveness), diameter.config snapshot (no YAML hot-reload), diameter.event_sink (file/none; clickhouse/kafka feature-gated). Inbound and outbound TCP+SCTP (peer::connect_with_transport). Failure answers are split by kind so the peer can act on them: 3002 only for a genuine no-route (no handler matched the application+command, or the matched handler returned None, which is the documented way to decline); 5012 DIAMETER_UNABLE_TO_COMPLY when siphon has a handler and could not carry it out — it raised, returned a non-DiameterAnswer, or produced an answer that would not serialize — logged at error with the handler's __qualname__ and the exception; 5014 for a message that did not parse. The split is load-bearing on Ro, where a 3002 to a CCR-UPDATE is read as a credit denial and tears the call down, so a script exception used to kill every live call at its first re-authorisation while looking like an OCS decision. Unit-tested in script::diameter_dispatch::handler_failure_tests (raise, async raise, wrong return type → 5012; None and no-handler → 3002; a working handler keeps its own code). Observability: siphon_diameter_inbound_requests_total{command}, siphon_diameter_inbound_answers_total{command,result_code} and siphon_diameter_inbound_duration_seconds{command} cover the served side — every other Diameter metric is client-side (what siphon sends, what came back), so a node in a server or DRA role previously carried none of its inbound load anywhere; alert on result_code="5012" for a script fault and "3002" for a routing gap. The ClickHouse/Kafka sinks are follow-ups. See examples/diameter_server.{py,yaml}. |
| AKA authentication (Milenage, local) | Implemented | auth.aka_credentials |
3GPP TS 35.206 — local key derivation without HSS; the 401 carries ck=/ik= for the P-CSCF's IPsec setup |
| AKA authentication (HSS-backed) | Production | auth.require_ims_digest() |
3GPP TS 33.203 via Cx MAR/MAA |
| IPsec SA management (P-CSCF) | Implemented | ipsec |
Shared protected client/server ports; SAs installed via direct XFRM netlink (Phase 3) with ip xfrm shell-out as fallback backend. The P-CSCF side of the SAs is a concrete UDP bind address: a family bound only to a wildcard gets no SA address (startup error, ipsec.allocate raises), because a 0.0.0.0 SA matches no inbound ESP (XfrmInNoStates). Protected REGISTER + de-REGISTER validated end-to-end in the sipp-ipsec harness, which asserts the SAs were activated |
| IPsec sec-agree primitives (script-driven) | Implemented | siphon.ipsec, request.parse_security_client(), reply.take_av(), Transform.alg / Transform.ealg |
3GPP TS 33.203 §6 + RFC 3329; Transform.alg/.ealg expose the ipsec-3gpp alg/ealg wire spellings (RFC 3329 Appendix A, TS 33.203 Annex H) without an allocated SA, so a script can build the Security-Server capability list that RFC 3329 §2.3.2 requires on a 421/494 (no authentication vector exists yet at that point, so SecurityServerParams is unreachable); HMAC-SHA-1-96 / HMAC-MD5-96 with NULL or AES-CBC-128 are the Annex H transforms, whose key expansion is Annex I (Rel-13 Annex H drops hmac-md5-96 and adds aes-gmac / aes-gcm, which siphon does not implement); HMAC-SHA-256-128 is a siphon extension, the RFC 4868 transform with a 256-bit key from siphon's own expansion HMAC-SHA-256(IK, "ipsec-int-sha256-128") (no 3GPP release checked here lists hmac-sha-256-128 in Annex H, and the label is arbitrary, so it interoperates only siphon-to-siphon); registration-tied lifetimes; IPv6; multi-instance SPI partitioning; multi-protocol XFRM selectors (TS 33.203 §6.3 has the SA pairs "all shared by TCP and UDP" and §7.1 says "The transport protocol selector shall allow UDP and TCP."; one SPI pair covers both ESP-over-UDP and ESP-over-TCP, required for iOS UEs mixing REGISTER/TCP with MO MESSAGE/UDP). siphon's Security-Server carries no transport protocol= parameter for any SA: no sec-agree spec defines one (RFC 3329 §2.2 and Appendix A, TS 33.203 Annex H, whose mech-parameters list is closed and whose own protocol rule is prot=ah|esp), and one pair carries both transports anyway. SecurityServerParams.protocol stays as an informational field naming what the pair is pinned to. A B2BUA INVITE that requires sec-agree is verified before the script runs (RFC 3329 §2.3.1): 494 unless it arrived over an active SA and its Security-Verify mirrors, parameter for parameter (q/prot/mod included; names, whitespace and parameter order not significant), the Security-Server recorded on the SA from the relayed REGISTER 401 (carried across PendingSA.refresh(), which keeps the SPIs; an SA with nothing recorded is checked on its own algorithms, SPIs and ports). A 494 over an SA carries its recorded Security-Server; a 494 to an unprotected request lists one mechanism-only ipsec-3gpp line per supported transform, a deliberate reading of RFC 3329 over TS 33.203 Annex H's mandatory spi/port. sec-agree counts as honoured for the Require check only on a verified call, and the caller's agreement (sec-agree in Require/Proxy-Require, Security-Verify, Security-Client) stays off the B-leg unless the script set it (ipsec::sec_agree, dispatcher::sec_agree_tests). |
| IPsec SA hard-lifetime repin on grant | Implemented | pending.activate(hard_lifetime_secs=…) |
XFRM_MSG_UPDSA on all four SAs; tightens kernel lifetime from the placeholder (UE's Expires ask, often 600000 s) to the registrar's grant on the 200 OK to auth REGISTER (3GPP TS 33.203 §7.4); kernel preserves add_time so deadline = original install + new value |
| IPsec SA hard-lifetime repin on REGISTER refresh | Implemented | automatic in registrar.save_proxy/save |
3GPP TS 33.203 §7.4: an IPsec-protected REGISTER refresh extends the bound SA pair's hard lifetime to the granted binding lifetime (granted Expires + 32 s Timer-F grace). IR.92 refreshes carry no AKA challenge (200-without-401 → no PendingSA → activate never fires), so this registrar hook is the only path that moves the SA forward on a refresh; without it an actively-refreshing UE's SA aged out at last-AKA + grace and was reaped + network-de-REGISTERed (live VoLTE/VoNR outage). The re-pin adds elapsed-since-install to the kernel hard_add_expires_seconds because XFRM_MSG_UPDSA preserves add_time (IpsecManager::update_sa_pair_lifetime keyed off SecurityAssociationPair::created_at), so the kernel deadline actually advances rather than staying pinned to the original install. Unit-tested (ipsec::tests elapsed-math + anchor-stability via mock kernel); needs root + live-core validation. |
| IPsec stale-pair cleanup on re-REGISTER | Implemented | pending.activate() (automatic) |
UE picks a fresh random port_uc on every REGISTER (TS 24.229 §5.1.1.2); without this, the manager's (ue_addr, port_uc)-keyed bookkeeping accumulated one entry per refresh and the prior pair's four XFRM policies leaked into the kernel forever. After enough cycles a new port_uc collided with a leaked selector and policy install hit EEXIST, breaking the registration. Activate now fire-and-forgets cleanup_other_pairs_for_ue to tear down every prior pair for the same UE address; the new pair (different port_uc by construction) installs cleanly. |
Protected re-/de-REGISTER without re-challenge (integrity-protected) |
Implemented | auth.stamp_integrity_protected() (P-CSCF), auth.verify_integrity_protected() (S-CSCF), contact.auth_user |
3GPP TS 24.229 trust chain for a REGISTER received over the SA. The P-CSCF stamps integrity-protected="yes" into each Authorization only when the request came over an SA negotiated for that header's IMPI (the SA records the IMPI of the REGISTER whose 401 keyed it), "no" otherwise, overwriting a UE-supplied value. The S-CSCF skips the AKA challenge only for yes/tls-yes/ip-assoc-yes from the IMPI recorded on a live binding of the To AoR (implicit set resolved); registrar.save() records request.auth_user per binding (persisted, legacy rows load as None, no carry-over). Both IMPI checks unit-tested; exercised end-to-end over real XFRM SAs in sipp-ipsec (de-REGISTER accepted without a challenge). Not yet validated against a live IMS core or handset. |
| Initial Filter Criteria (iFC) | Production | isc |
XML trigger-point matching + per-user profile storage from Cx SAR |
| IMS P-CSCF role | Production | Example examples/ims_pcscf.{py,yaml} |
|
| IMS I-CSCF role | Production | Example examples/ims_icscf.{py,yaml} |
|
| IMS S-CSCF role | Production | Example examples/ims_scscf.{py,yaml} |
|
| 5G SBI — Npcf (policy) | Implemented | sbi |
N5 app-session for VoNR QoS. sbi.create_session(media_components=[…]) builds the spec-correct TS 29.514 AppSessionContext: request data nested under ascReqData (a flat body left the PCF reading ueIpv4 as null → session created but never bound), medComponents/medSubComps as maps keyed by medCompN/fNum (not arrays) with the exact wire names medCompN/medType/fStatus/codecs/fDescs/flowUsage and hyphenated ENABLED-UPLINK/ENABLED-DOWNLINK so PCF gating works on real UPFs; same dict shape as diameter.rx_aar. The created appSessionId is taken from the 201 Location header (it is not a body field); modify is an application/merge-patch+json PATCH of AppSessionContextUpdateDataPatch, with the update data under ascReqData like create (it was sent flat, where a PCF following the OpenAPI finds nothing to modify; body-capture tested). Per-call pcf_uri= addresses a session at a discovered PCF instead of the static npcf_url; create returns app_session_uri and update/delete accept it for replica-independent teardown. Wire format corrected after a live PCF trace exposed the missing envelope; message-level + SDK tested (axum body-capture asserts the ascReqData envelope and medComponents map on the wire), live re-validation against a PCF pending. Inbound PCF event notifications (@sbi.on_event, TS 29.514 EventsNotification) are now passed to the script verbatim as a dict — previously they were projected through a lossy typed struct that dropped the required evSubsUri correlation key and 422'd (silently lost) any notification carrying flows (the spec shape is {medCompN, fNums}, not {flowId}). Both TS 29.514 AF callbacks are served on sbi.notif_listen: POST /sbi/events/notify (EventsNotification → @sbi.on_event) and POST /sbi/events/terminate (TerminationInfo → @sbi.on_terminate, the N5 counterpart of an Rx ASR). Only the bare /sbi/events was served before, so both TS 29.514 callback paths 404'd and a PCF termination never reached the script; the bare route still dispatches to on_event and is deprecated. Answers: 204 once the handlers ran (a raising handler is logged), 400 non-JSON, 503 when the Python executor sheds the job (was a 204 for work no handler saw). Router tested with oneshot against real Python handlers; not yet validated against a live PCF. create_session(events=[…], notif_uri=…) subscribes to PCF events: ascReqData.evSubsc is sent as EventsSubscReqData (events of {event, notifMethod: EVENT_DETECTION} plus notifUri); before, no subscription was ever sent, so on_event had nothing to receive. update_session(events=…, notif_uri=…) replaces it in the merge patch and never sends null. Body-capture tested on create and modify. |
| 5G SBI — Nbsf (PCF discovery) | Implemented | sbi.discover_pcf_binding, sbi.bsf_url |
Nbsf_Management pcfBindings lookup keyed on the UE IP (TS 29.521) — the reliable 5G-vs-4G discriminator a P-CSCF uses to pick N5 vs Rx per session. 200→binding dict (incl. ready-to-use pcf_uri), 404→None (4G), 5xx/timeout→sbi.BsfError. Message-level + SDK tested (axum mock); not yet validated against a live open5gs BSF. |
| 5G SBI — SCP indirect communication | Implemented | sbi.communication: indirect |
Spec-compliant indirect routing via the SCP (TS 29.500 §6.10). Npcf Model C emits 3gpp-Sbi-Target-apiRoot (the PCF known from the BSF binding); Nbsf Model D (delegated discovery) emits 3gpp-Sbi-Discovery-target-nf-type: BSF / service-names: nbsf-management / requester-nf-type (default AF). direct (default) is byte-identical to today. Header-level tested (axum mock); not yet validated against a live SCP. |
| 5G SBI — Nchf (charging) | Implemented | sbi |
External Control Plane¶
An out-of-process application drives B2BUA calls over a WebSocket, in the model Asterisk
has with ARI and FreeSWITCH with ESL. A Python handler hands a call over with
call.handover("app"); siphon holds the INVITE un-dialed and emits StasisStart, and the
application answers, routes, transfers or hangs up over the socket.
Every verb's reply reports the local action only. A far-end outcome — the callee answering, a transfer completing — always arrives later as an event, never folded into a command reply.
| Feature | Readiness | Config | Notes |
|---|---|---|---|
Call handover (call.handover) |
Implemented | control.apps[] |
siphon holds the INVITE un-dialed and emits StasisStart with the full SIP context (real headers, source, R-URI shape, body) plus a stable {channel, call_id, sip_call_id} id triple that joins CDR and HEP with no mapping table. Answer-first mode (answer=True, ws_uri=…) answers and anchors media to the WebSocket bridge before handover, so the app drives an already-connected channel; on a backend that cannot do it the handover fails visibly rather than returning a fake 200 |
| Inbound WebSocket listener | Implemented | control.listen |
Persistent connection; per-app bearer tokens, constant-time compared, feeding the auto-ban store |
| Outbound per-call connect | Implemented | control.apps[].per_call_connect + control.apps[].connect_url |
siphon dials the controller at handover, so the instance that accepts the connection owns the call. The multi-instance default: no distributed lock, no claim key, and pairing the per-call ws_uri with it puts the audio socket on the same instance by construction |
| Exactly-one-owner dispatch | Implemented | — | Per-tenant scoping; a cross-app target answers forbidden |
| Handoff deadline | Implemented | control.limits.handoff_deadline_ms |
A safe default action (503) when no controller acts in time, so a call is never left parked on an absent app. control.inbound.deadline_ms overrides it for the script-free inbound path |
resync reattach |
Implemented | — | A controller that reconnects inside the grace window recovers its calls instead of losing them |
| Backpressure | Implemented | — | Bounded per-connection outbound queue with per-call event/reply ordering and drop-oldest overflow; a slow controller can never stall the datapath |
Call verbs: answer ring progress reject hangup drop route |
Implemented | — | Alerting and early media are two verbs, not one: ring sends a plain 180 Ringing (RFC 3261 §13.2.1 — alerting, no session semantics) and refuses a body, since SDP on an 18x is early media (RFC 3960 §3.1); progress is the one that opens an early-media path and still takes any 1xx, defaulting to 183 Session Progress. That split is what lets an application ring for an interval of its own policy's choosing before answering. A provisional's reply names which it was — {state: "ringing"\|"progress", code, early_media} — in the same vocabulary as the callee-side ChannelStateChange. drop is the third teardown and the only silent one: it abandons an unanswered call with nothing on the wire — no final response, no CANCEL — and releases it, so an unsolicited INVITE costs an enumeration sweep silence instead of a 404 that confirms the number it probed. It CANCELs any B-leg a dial left ringing (RFC 3261 §9.1), stops the reliable provisionals to the caller (RFC 3262 §3), deletes anchored media and removes the actor; an answered call is refused invalid_state because its dialog is owed a BYE (§15). The reason reaches the log and the CDR (disconnect_initiator: "control", sip_reason, response_code: 0) so a dropped call never reads as a leak |
originate |
Implemented | — | A call siphon places itself — the primitive under click-to-dial, callbacks and the dial half of a transfer. The channel id is supplied by the caller, never minted, so an application stages per-call context before anything reaches the network; a duplicate live id answers the distinct conflict code. Asynchronous: the reply is the local action, ringing and answer follow as events. ACKs the 2xx (RFC 3261 §13.2.2.4), ACKs a final non-2xx on the INVITE's own branch (§17.1.1.3), and CANCELs rather than answering when abandoned before answer (§9.1). Media requires exactly one plan — a verbatim sdp, or media: true for an offerless INVITE answered locally — because an unanswerable 2xx is a connected call with no audio. originate {aor} rings every phone registered at an AoR, each as its own originated call over its own flow and Path (a phone on TCP, TLS or WSS behind NAT is reachable no other way), parallel or sequential under a group deadline; the first to answer becomes the channel's call and the rest are CANCELled, a late 2xx ACKed with every stream rejected and BYEd. Dispatcher-driven tests in dispatcher::originate_group_tests (parallel win, sequential advance on 486 and on ring timeout, all-fail, cancel, deadline, TCP flow, Path, wildcard listener, and a drain-to-zero leak gate on the group store) and dispatcher::control_originate_aor_tests (the verb end to end). Not yet SIPp-validated |
dial target to (called party) |
Implemented | control dial |
A target names the B-leg's To URI with to, on {uri} and {aor} targets, per branch, for parallel, sequential and bridge dials. Without it To keeps the caller's user at the target's host, which on a divert addresses the B-leg to the originally dialled number. Set in the B-leg builder before the dialog state is captured, so siphon's later in-dialog requests carry it (RFC 3261 §12.2.1.1). A non-SIP-URI to is refused before anything rings. The per-target identity fields (from, from_display, p_asserted_identity, privacy) apply to {aor} targets too, on every contact (they were dropped on that form before; control::sip_adapter parser test and dispatcher::control_dial_bridge_args_tests). Dispatcher-driven tests in dispatcher::control_dial_to_tests (UDP egress and the leg's recorded dialog To) and dispatcher::control_dial_bridge_args_tests; not SIPp-validated. |
dial {on_answer: "bridge"} |
Implemented | control dial |
Ring phones for a caller the controller already answered and anchored (the end of an IVR flow) and bridge the one that picks up. Each phone is its own originated leg (an originate group: {aor} contacts over their own flow and Path, URIs as written; parallel / sequential), presenting the dial's identity arguments or the caller's From; the caller's own dialog and the connecting dial's B-leg machinery are never touched. Ringback (ringback, default ringback_eu, false for none) starts on the first 180-183 (RFC 3960), through the play {tone} engine path, held behind a prompt still playing until its PlayFinished; stopped at the bridge and before DialFailed; its play events carry origin: "ringback". Answers are provisional (an originate group whose answers are confirmed): the other phones ring on until the bridge path has joined the answering phone to the caller, a phone answering meanwhile waits as an ACKed standby, and a bridge that fails hangs its phone up (Q.850 41, DialBranchFailed cause bridge_failed) and falls back to the standbys in answer order, the phones still ringing, or a sequential dial's next target. Only the bridged phone gets DialAnswered, with a minted channel (same app, connection and on_lost as the caller's). A caller that hangs up while the phones ring CANCELs every one. Phone early media is not relayed. The dial's profile describes the phone: the bridge offers the phone with it and re-INVITEs the caller with the caller's own, and a failed bridge leaves the caller's media session intact so the ringback resumes on it (dispatcher::control_bridge_media_tests, against a native engine mock that refuses a command on a call it no longer holds). Whether each party's media ingress is pinned to its signalling source is that party's own profile's decision, so a phone behind NAT rung with a profile that asks for it is pinned whatever the caller was answered with (dispatcher::control_bridge_ingress_tests). Dispatcher-driven tests in dispatcher::control_dial_bridge_tests (parallel win to a formed bridge, sequential advance on 486 and ring timeout, all-fail, caller BYE and siphon teardown while ringing, bridge refused by the phone and unable to start, and a drain-to-baseline leak gate on the dial, group and call stores), dispatcher::control_dial_bridge_fallback_tests (a ringing phone and a standby bridged after a failed bridge, a standby released when the bridge forms, sequential advance after a failed bridge) and dispatcher::control_dial_bridge_args_tests (refusals, identity, ringback). SIPp: the sipp-dial-bridge job (run-tests.sh --dial-bridge) |
cancel_dial |
Implemented | Unit + SIPp (--control-transfer) |
Give up on a ringing dial and keep the caller: the phones are CANCELled (RFC 3261 §9.1) and the dial ends in DialFailed 487, for a bridging dial with the controller's reason as cause. The caller stays as the dial found it, owned and free to be dialled again. Refused invalid_state when nothing rings and once a phone is being bridged. route is refused while a dial rings. The reply is held until the dial has let go of the caller, so DialFailed is queued ahead of it and a dial or route sent on reading it is taken (control_cancel_dial_tests::a_cancel_replies_once_the_dial_has_let_go_of_the_caller; the watch a cancel waits on and the per-dial release in dispatcher::b2bua::dial_bridge a_dials_conclusion_closes_when_it_lets_go_and_never_takes_another_dials_claim). Unit-tested end to end through the command consumer in control_cancel_dial_tests and, for a connecting dial, a_cancelled_connecting_dial_cancels_every_branch_and_parks_the_caller. SIPp: the cancel-dial case of the sipp-control-transfer job cancels a bridging dial that rings two phones, one alerting and one that only answered 100: both are CANCELled and their 487 ACKed, the caller receives nothing and answers an in-dialog INFO afterwards, DialBranchFailed for each phone precedes DialFailed 487 with the controller's reason, a second cancel and a route mid-ring are refused, and a second dial for the same caller rings on new dialogs. A phone that has sent no response when the dial is cancelled is not CANCELled until its first provisional (RFC 3261 §9.1): unit-tested in originated_cancel_awaits_provisional_tests (a late 180 and a late 100 Trying each draw the CANCEL, a late 200 is ACKed and BYEd with no CANCEL, a late failure is ACKed only, and a phone that never responds ends at Timer B with no CANCEL, a 2xx after it still ACKed and BYEd, and nothing kept past the branch's expiry), and on the wire by the cancel-late-ringing and cancel-late-answer cases of the same job. A connecting dial's cancel is unit-tested only. |
bridge / unbridge |
Implemented | Unit + SIPp (--bridge) |
Joining two legs the process already owns: 3PCC re-negotiation across two answered dialogs (both re-offered, peer first), a media re-anchor across two call actors, RFC 3261 §14.1 glare refused rather than guessed, and a peer-hangup policy. unbridge leaves both legs answered and held (RFC 3264 §8.4). Each leg is offered what its own profile describes, or one pair profile names; both legs' own media sessions are deleted only once the bridge forms, so a refused bridge leaves both usable (dispatcher::control_bridge_media_tests). The received_from source hint follows the profile of the party whose SDP each engine command carries, not the profile that shapes the command (the other party's), so a party behind NAT is pinned to its signalling source by its own profile whatever the party it is joined to was anchored with, on the bridge's offer and answer and on a relayed re-offer; a pair profile decides for both (dispatcher::control_bridge_ingress_tests, read off the commands an in-process native engine records, since a SIPp party sends media from the address it signals). A re-INVITE or UPDATE with SDP on either leg of a formed bridge is relayed to the other leg through the pair's media session, each party shaped by its own side of the bridge; refusals relayed back with the session restored, crossing offers 491, bodyless refreshes answered from the session in force, 487 / 408 for a hangup or no answer mid-relay (dispatcher::control_bridge_relay_tests, leak gate in dispatcher::b2bua::bridge_relay). SIPp --dial-bridge (the sipp-dial-bridge job) holds and resumes the caller across the formed bridge. The SIPp mode asserts audio flowed both ways via the engine's per-leg counters, not just that the verb returned ok. A pair bridged again after an unbridge is shaped and pinned as its first bridge was: the with leg's own session was retired when that bridge formed, so its profile and its received_from policy are read off the pair's session, in either order the two legs are named, and never applied to a different leg bridged to the anchor afterwards (dispatcher::control_rebridge_tests, with a drain of the session store and the engine; b2bua::bridge a_pair_bridged_again_reads_the_peers_own_profile_off_the_pairs_session). A with leg bridged to a different anchor after an unbridge is offered and pinned as its own profile describes: what the leg was anchored with is kept per call beside the media sessions (rtpengine::session::own_media) and read by the next bridge, in either order the legs are named; the record goes with the call on a BYE from the leg, a BYE from the anchor it is bridged to, a teardown of siphon's own and a BYE while parted (dispatcher::control_bridge_elsewhere_tests, including the drain-to-zero gate the_per_call_media_record_drains_on_every_way_a_call_ends). The record is what the call was first anchored with, written once for the with leg and for the anchor alike and never rewritten by a later bridge, so a pair profile describes only the pair it was named for: an anchor bridged under one, parted and bridged to another leg with no profile is re-INVITEd with its own transport and pinned by its own policy (an_anchor_bridged_under_a_pair_profile_is_its_own_again_for_the_next_bridge; rtpengine::session::own_media unit tests; the drain gate counts both anchors' records too). A leg parted by unbridge has a re-INVITE or UPDATE answered from its own dialog, with nothing sent to the engine or the other leg: a request that changes nothing (no SDP, or an unchanged o=) 200 with the session in force, an offer that would change the session 488, after which the pair bridges again on the engine call it still holds (dispatcher::control_unbridge_reoffer_tests). Limitation: a parted leg cannot change its own media: its own hold or resume, a new address or a new codec list is refused 488 until it is bridged again, and siphon does not follow a parted leg that moved. The application bridges the leg again before it re-offers; the endpoint keeps its session and may retry after the 488 (RFC 3261 §14.1), and a re-offer after the bridge forms is relayed. Not SIPp-covered: a parted leg's re-offer, and a bridge to a different anchor |
Media verbs: play stop dtmf hold unhold |
Implemented | — | Bound to the configured media backend, with typed replies rather than a hang: no anchored session → not_found, backend cannot → unsupported_verb, otherwise unavailable. hold is implemented as silence; a gate verb is a follow-up play's repeat is a total play count or "inf" to play until stopped (siphon-rtp; a count-only backend refuses it), and an argument of the wrong type is bad_request instead of being read as absent. hold silences the call's media in both directions and sends no SIP; on a call the engine only relays it answers invalid_state (media_not_processed). Unit-tested in play_repeat_is_a_count_or_inf, play_refuses_an_argument_it_cannot_use_rather_than_dropping_it, an_endless_play_reaches_the_engine_as_inf, hold_on_a_call_the_engine_only_relays_is_an_invalid_state. SIPp: play / stop against a mock engine in the sipp-control job, and against the real siphon-rtp engine in the media case of the sipp-control-transfer job, where hold on a two-party call relayed on one codec is refused media_not_processed by the engine's own answer, play {repeat: "forever"} is bad_request naming repeat, and play {repeat: "inf"} plays until stop ends it (PlayFinished, completed: false). dtmf, unhold and a hold that succeeds are not SIPp-covered. |
stream_start / stream_stop |
Implemented | media.backend: siphon-rtp |
Attach and detach a WebSocket audio stream mid-call. mode: tee (the default) streams a copy while the call keeps relaying; mode: bridge is a takeover — the WebSocket server becomes the leg's far side and A↔B is unwired — and a stream_start on a call that already has a bridge re-points it in place rather than failing, so a party can move between media servers without the gap a detach-then-attach would leave. A bridge stop is refused where there is no relay to hand the call back to (a ws_uri-negotiated bridge, or a single-leg takeover), unlike the idempotent tee stop. Native backend only — rtpengine and rtpproxy raise a typed error rather than reporting success while streaming nothing A bridge takes profile (a media profile name) for its wire rate, noise suppression, echo cancellation, VAD and barge-in, read as answer with ws_uri reads them, and so does rtpengine.attach_ws_bridge(profile=); unknown names refused before anything is sent. Needs siphon-rtp 0.11+. Unit-tested on the wire (siphon_rtp::tests::attach_ws_bridge_*, answer_with_sdp::attach_ws_bridge_*, sip_adapter::tests::stream_start_*profile*); not yet validated against a live engine. |
DialBranch / DialBranchFailed / DialAnswered |
Implemented | — | Every B-leg a control-plane dial rings is named to the controller by leg_id and leg_sip_call_id (the Call-ID its INVITE carries), at creation (each fork branch, and each attempt of a sequential hunt when the failover engine places it), and again with its outcome: rejected / timeout / cancelled / unsent, or DialAnswered for the winner. DialFailed lists every branch with its outcome. Each branch is reported once. Dispatcher-driven tests in dispatcher::control_dial_branch_events_tests assert each event against the Call-ID of the INVITE read off the UDP egress. Not SIPp-validated |
DialogStateChanged (application-level, events: [dialog]) |
Implemented | control.apps[].events: [dialog], control.dialog_state |
RFC 4235 dialog state of every dialog a registered AoR has through siphon, B2BUA and proxy alike, for a controller serving the dialog event package (busy-lamp field): trying / proceeding / early / confirmed / terminated, direction initiator/recipient, the phone's own Call-ID and tags, and the remote identity. A call a phone places is matched by its From AoR only when a live binding vouches for the INVITE (its authenticated identity, or the address the binding registered from); a callee by the {aor} target it came from, the unique binding whose Contact is the Request-URI, or the binding its captured flow names. B2BUA: the A-leg, every B-leg (script dial/fork, control dial), transfers (terminate, transparent, siphon-originated REFER), a Replaces takeover, originate. Proxy: every relayed or forked INVITE with a registered party, tracked per branch by dialog (Call-ID + tags) and Record-Routed by siphon so the BYE from either end comes back (the entry is marked dlgw and its in-dialog requests are routed by siphon without the script); a BYE the script answers itself still ends it. terminated is reported once on every teardown: BYE either side, CANCEL, final failure, ring timeout, a branch cancelled because another answered, an answer-glare loser (never shown confirmed), a referrer, a replaced party, every call removal. For a proxied dialog the reported state (never the call) also ends when the phone's binding is removed/expires/is reaped, a negotiated RFC 4028 interval runs out, an in-dialog OPTIONS probe is answered 481 or goes unanswered probe_failures times, it rings past max_early_secs, or outlives max_lifetime_secs; a B2BUA leg ends on the binding check too. Branch events of a control dial also carry the aor an {aor} target resolved to. Tests: dispatcher::dialog_state_events_tests (B2BUA ring group, phone-placed call, spoofed From, CANCEL, failed branch, ring timeout, answer glare, originate, binding removal, drain), dispatcher::dialog_state_transfer_tests (terminate / transparent / siphon-originated REFER, Replaces), dispatcher::proxy_dialog_state_tests (proxied ring group with forced Record-Route and BYE, script-answered BYE, in-dialog routing past a script that cannot route, CANCEL, declining branch, untracked INVITE left alone, probe 481/timeout/answered, session interval, binding removal, Timer C and lifetime bounds, drain), proxy::dialog_state::tests; every state sequence and Call-ID/tag checked against the egress. Cost: benches/dialog_state.rs. Not SIPp-validated |
PlayStarted / PlayFinished |
Implemented | media.backend: siphon-rtp |
Both halves of a playback's lifecycle, on the control rail and as @rtpengine.on_play_finished in-process. play over the rail is always fire-and-forget, so PlayFinished — not the accept's estimated duration — is when the prompt is actually over. Carries the correlating play_id, the end reason and a completed flag, since a stop, a supersede and a decode error all end a playback without it having been heard in full. A blocking in-process play_media still returns its own outcome; the event fires regardless |
WsTeeStarted / WsTeeEnded |
Implemented | media.backend: siphon-rtp |
The lifecycle of a mode: tee stream, on the control rail and as @rtpengine.on_ws_tee_started / on_ws_tee_ended in-process. Exactly one end per start, including when the server ends it. frames_dropped > 0 means the consumer could not keep up — the call itself is never affected by a tee. detached and call_ended are orderly (unexpected: false); server_closed, server_stopped and transport_error are not |
WsBridgeStarted / WsBridgeEnded |
Implemented | media.backend: siphon-rtp |
The lifecycle of a mode: bridge stream, on the control rail and as @rtpengine.on_ws_bridge_started / on_ws_bridge_ended in-process. Exactly one end per start, including when the server ends it; a re-point is an ended+started pair. detached and call_ended are orderly (unexpected: false) — every other reason leaves a live call with no media far side, so an unexpected end is logged at WARN even with no handler registered |
MediaSummary (control rail) |
Implemented | media.backend: siphon-rtp |
The engine's summary of a media session it ended (per-leg counters, loss, jitter, RTT, MOS where measured) on the channel owning its SIP Call-ID, alongside the media CDR. On an ordinary hang-up it arrives after StasisEnd, routed by a 30 s tombstone (SIP Call-ID → owning connection + channel id) that is spent on delivery, expires, or goes with the owner's connection; the only event that can follow StasisEnd. A bridged pair's summary, which names the pair's own engine call, reaches both legs' channels once each: the media store keeps engine call → SIP Call-IDs while the session lives and for 30 s after. Its media CDR is written once per leg on that leg's Call-ID, each carrying media_call_id + media_parties so a collector counts the pair once. Leak-gated by control::registry::tombstone_tests and rtpengine::session::session_parties_tests. Each leg's media_started_at_unix_ms is carried when the engine reports one |
MediaStarted (control rail) + @rtpengine.on_media_started |
Implemented | media.backend: siphon-rtp 0.10+ |
The engine's first-packet report per leg, once per leg, on the channel of the party the leg faces (a bridged pair routes by the leg's tag through the media store's engine-call parties). Payload carries the leg, both tags, the latched source, the signalled address and nat_rewritten. Stateless in siphon: nothing is kept per event. Each offer/answer also names the leg's SIP Call-ID to the engine for HEP correlation |
replace_peer |
Implemented | Unit + SIPp (--control-transfer) |
Swap one party of an answered call for a freshly dialed target, with no REFER anywhere — the same dial / promote / BYE round a siphon-terminated REFER runs, on the app's say-so. The replaced leg stays up while the target rings and is released only once it answers, so the survivor hears ringback rather than dead air and a refusal leaves the call untouched; the alternative a script would otherwise hand-roll (hang up, then re-INVITE) also skips on_bye, the CDR, the charging stop and the media release. Outcome arrives as PeerReplaced / ReplaceFailed, never in the reply. In-process twin: b2bua.replace_peer() target may be {aor} (dialled over the registered contact's flow and Path; an AoR with several registered contacts rings them all, keeps the first to answer and CANCELs the rest, each loser's 487 ACKed and a 2xx that crosses its CANCEL ACKed and released with a BYE), and from / from_display / p_asserted_identity / privacy / headers shape the new leg. Unit-tested in a_replacement_leg_is_called_as_its_aor_and_presents_the_named_identity, a_transfer_target_is_a_uri_or_a_registered_aor, a_transfer_to_an_aor_with_several_contacts_rings_every_one, and replacement_fork_tests (the first answer wins under two concurrent 2xx, a late 2xx is released, both failing reports one failure with the RFC 3261 §16.7 best status, the deadline and a hangup CANCEL every ringing contact, and each unanswered contact's engine call is deleted). SIPp: b2bua-replace-peer (in-process, one target) in the sipp-b2bua-refer job, and the replace-aor / replace-aor-refused cases of the sipp-control-transfer job, which replace the callee with an AoR two contacts registered at. In the first both are rung on Call-IDs of their own, one answers and is ACKed, the other is CANCELled and its 487 ACKed, the surviving party is re-INVITEd, the replaced one BYEd, and the controller gets one PeerReplaced naming the answering contact's dialog. In the second the contacts answer 486 and, a second later, 480: one ReplaceFailed with status 486 and the call kept, shown by an INFO the caller sends through to the phone it was connected to. Not SIPp-covered: a 2xx crossing its CANCEL, replace_a_leg, and a replacement on a media-anchored call. On a media-anchored call the fresh engine call's offer and answer each carry the received_from hint of the party whose SDP they hold, by that party's own policy: a profile named for the pair decides for both (offer half the surviving party, answer half the target), an inherited one leaves the surviving party the half it was set up under and gives the target the replaced party's; each of several ringing targets is pinned by its own source, and a second replacement of the same call reads the same policies again (dispatcher::transfer_ingress_tests, read off the commands an in-process native engine records, since a SIPp party sends media from the address it signals). Replacing the callee keeps the call's media session under the caller's Call-ID, naming the caller first and the target second, and confirms the target's leg when its 2xx is ACKed: a hold from the caller afterwards is re-offered on the pair's engine call under the caller's own tag and relayed to the target instead of being refused 491, and the pair's engine call is deleted with the call (dispatcher::replacement_fork_tests a_replaced_callee_leaves_the_calls_media_session_on_the_new_pair). A target no INVITE can be sent to is refused with no replacement left on the call, so the next one is dialled; for an accepted REFER the subscription is ended with a 503 sipfrag NOTIFY and ReplaceFailed, and of several targets the ones that were dialled ring on (dispatcher::undialled_replacement_tests, a_replacement_with_no_target_entered_is_failed_and_one_with_a_target_is_not). A re-INVITE or UPDATE from either party after a Replaces takeover is pinned by what the re-paired session recorded, the surviving caller its offer half and the newcomer the replaced callee's answer half, though the takeover moved them to the other slot of the call (dispatcher::reoffer_ingress_tests a_reoffer_after_a_takeover_pins_each_party_by_what_the_pair_recorded) |
| Replacement / transfer target timeout | Implemented | — | A replacement runs on an answered call, which the answer-timeout sweep deliberately skips, so a target that sends a 180 and then nothing left the transfer armed for the life of the call and the survivor bridged to nobody. A per-replacement deadline now cancels the target and restores the original pair; it covers REFER-driven transfers too, which had the same hole |
Header verbs: set_header get_header remove_header |
Implemented | — | In-process scripts additionally have get_headers(name) for a header the peer spread over several lines (RFC 3261 §7.3.1) — get_header reads only the first |
Inbound REFER → TransferRequested |
Implemented | Unit + SIPp (--control-transfer) |
A REFER on a controlled call is held un-answered and the decision goes to the owning app via accept_refer / reject_refer. If the app never decides, a deadline sweep answers 603 Decline so a REFER is never left pending. An uncontrolled call still runs the in-process handler A REFER from the B-leg is resolved through the call and reaches the owning channel (referrer_leg, referrer_sip_call_id); a second REFER while one is undecided is answered 491, and so is a REFER from either party while a transfer is still being carried out on the call (a terminate transfer or replace_peer whose target rings, a transparent REFER the far end has not answered), with no event and no @b2bua.on_refer, until that transfer succeeds, fails or times out (dispatcher::refer_answer_tests: a_refer_is_refused_while_a_transfer_is_in_flight_and_taken_once_it_fails, …_runs_out_of_time, …_has_succeeded, a_script_is_not_shown_a_refer_while_its_transfer_is_in_flight, a_refer_is_refused_while_a_relayed_refer_awaits_the_far_end, a_refer_is_refused_while_a_replace_peer_is_in_flight). An attended transfer naming a dialog this node hosts reports it as replaces.local. accept_refer takes the {aor} target and identity arguments replace_peer does. accept_refer {mode: "controller"} answers 202, sends the first sipfrag NOTIFY and dials nothing; the application carries the transfer out and reports with complete_refer {code, reason?}, which sends the terminating NOTIFY. A deadline (timeout, default 60 s, at most 180) reports 503 for an application that does not and raises TransferTimedOut ({reason: "timeout", code, referrer_leg}) on its channel, once (an_application_that_misses_its_deadline_is_told, through the deadline sweep to the application's event queue); the referrer's BYE ends the subscription silently; a further REFER meanwhile is answered 491. Not implemented: joining two hosted calls locally for an attended transfer, and a siphon-terminated transfer on a call bridged from two legs (the controller mode is how an application does both itself). Unit-tested in a_refer_from_the_callee_of_a_controlled_call_reaches_its_application, an_attended_refer_names_the_hosted_dialog_it_replaces and, for the controller mode, control_refer_controller_tests (through the command consumer on a test dispatcher). SIPp, both in the sipp-control-transfer job with the REFER sent by the callee: refer-callee (TransferRequested names leg b and the Call-ID of the INVITE that phone received; reject_refer reaches the referrer as the final response named; accept_refer in terminate mode dials the target, sends the referrer 202 and the sipfrag NOTIFYs ending in 200 OK / terminated, BYEs it and reports PeerReplaced) and refer-controller (accept_refer {mode: "controller"} sends 202 and a 100 Trying NOTIFY with active;expires=<timeout> and never dials the Refer-To target; complete_refer sends the terminating sipfrag for a 200 and for a 486 with its reason phrase, a second report is refused, and a transfer left unreported ends in a 503 sipfrag at its deadline). Not SIPp-covered on the control rail: a REFER from the caller's leg, transparent mode, an attended transfer (replaces.local), the 491 for a second REFER and the decision deadline's 603. A siphon-terminated REFER and an INVITE with Replaces on a media-anchored call pin each party's media ingress to its signalling source by its own policy on the fresh engine call, from either side of the call (dispatcher::transfer_ingress_tests). A REFER already decided on is recognised when retransmitted (Call-ID, CSeq and Via branch) and gets its final response again for 64·T1, with no second TransferRequested, dial or relay, in terminate, transparent (absorbed until the far end answers), after a rejection and after the deadline's 603, and on a call its script accepts (dispatcher::refer_answer_tests; store leak gate answered_refer_store_drains_to_baseline). A REFER still undecided when its call ends is answered then and released: 487 on its sender's own BYE, 603 on the other party's BYE or a teardown, 481 from an accept that finds the call gone (dispatcher::refer_answer_tests). A held REFER is kept under the call, not under the A-leg Call-ID its channel was bound to, and its sender is recognised by the REFER's own dialog: after a Replaces takeover it is still answered when the call ends, still decided through the channel, follows a caller the takeover moved to the other leg, and is answered 487 when the party taken over is its sender (dispatcher::refer_answer_tests: a_refer_held_across_a_takeover_of_the_caller_is_answered_when_the_call_ends, a_refer_held_across_a_takeover_is_still_decided_by_its_channel, a_refer_follows_its_sender_when_a_takeover_moves_it_to_the_other_leg, a_refer_accepted_after_a_takeover_names_its_sender_as_the_referrer, a_referrer_that_is_taken_over_has_its_held_refer_answered; store exits and drain in pending_inbound_refer_is_taken_by_its_dialog_or_its_channel_and_drains) |
| Outbound REFER verdict | Implemented | — | TransferProgress while it moves, then exactly one TransferCompleted / TransferFailed. RFC 3515 §2.4.4 splits "accepted for processing" (the 2xx to the REFER) from the real outcome (the message/sipfrag NOTIFY), so a 2xx is progress and never completion. The 1-based attempt distinguishes a carrier that challenged and was answered from one that refused, when both arrive as 407. A terminating NOTIFY always yields a terminal stage — including when the body carries no readable status — and a call torn down mid-transfer flushes a call_ended failure before StasisEnd |
Events: StasisStart StasisEnd ChannelStateChange ChannelDtmfReceived PlayStarted |
Implemented | — | StasisEnd always carries the hangup cause, and the SIP status (code + response) on every teardown that had a final response: 487 on a CANCEL in either direction (RFC 3261 §9.1/§9.2), 408 on the answer timeout, the callee's own status on a rejected originated leg, and the status siphon sent on a reject / an unanswered hangup / the handoff-deadline default. A teardown with no SIP response (a plain BYE) omits both rather than inventing one, so code present always means a status was on the wire. PlayStarted reports a play the backend accepted, carrying the play_id a targeted stop addresses — the media contract is accept-on-start, so it means the engine armed the playback, not that audio has reached the wire; a refused play emits nothing, so "no start event yet" reads as "not started". On a bridged pair, which relays both legs through one engine call, ChannelDtmfReceived and the per-party media events go to the leg whose engine tag they carry, never both (dispatcher::bridged_pair_media_events_tests) |
| Adapter API for extensions | Implemented | SiphonServer::register_control_adapter |
ControlAdapter trait with an opaque JSON DTO seam, so a protocol extension registers its own control surface on the same rail. The built-in SIP adapter ships in core |
| Client SDKs | Implemented | — | siphon-control (PyPI), siphon-control-client (crates.io) and the TypeScript client. They version independently of siphon core against the protocol, on their own release train |
| Prometheus metrics | Implemented | metrics.prometheus |
siphon_control_connections, siphon_control_controlled_calls, siphon_control_commands_total, siphon_control_events_dropped_total, siphon_control_auth_failures_total, siphon_control_handoff_timeouts_total |
| Refusal logging + typed refusal details | Implemented | — | Every applied command logs one line: control plane: command applied at debug, control plane: command refused at warn (the controller asked for something impossible) or error (unavailable — the stack could not do something possible), carrying app, module, verb, channel, sip_call_id and, on a refusal, the stable code and the message the controller got. A failed reply may carry error.details, machine-readable fields beside the prose ({verb, argument, bytes, limit_bytes} on an oversized play blob), so a controller branches without parsing English. Previously a refused verb wrote nothing at any level and an oversized inline blob came back as unavailable from the frame encoder, naming a JSON frame length rather than the argument |
| Functional (SIPp) coverage | Planned | — | Five jobs run on every pull request: sipp-originate (originate), sipp-control (handover, handover from on_failure, the handoff deadline, media and recording verbs against a mock engine, early media, hold and resume, in-dialog INFO, a dial that fails, the script-free control.inbound path, one-owner dispatch, resync), sipp-bridge (bridge / unbridge with audio counted on a real engine), sipp-dial-bridge (dial {on_answer: "bridge"} on a real engine: a registered phone rung offerless for an answered caller, ringback, the bridge, the caller's hold and resume relayed, and the phone released when the caller leaves) and sipp-control-transfer (cancel_dial, a REFER from the callee through reject_refer, accept_refer in terminate and controller mode and complete_refer, replace_peer to an AoR with two contacts, and hold / play {repeat} on a real engine). The same job runs cancel-late-ringing and cancel-late-answer, which give a bridging dial up before its phone has sent any response: the phone receives nothing but its INVITE's retransmissions until it speaks, its late 180 draws the CANCEL and the ACK of its 487, and its late 200 is ACKed with an answer and released with a BYE without ever being CANCELled (RFC 3261 §9.1). Still unit-tested only: originate {aor}, drop, route, refer and its verdict events, stream_start / stream_stop, the header verbs and DialogStateChanged |
Lawful Intercept / Recording¶
| Feature | Readiness | Config | Notes |
|---|---|---|---|
| LI master switch + audit log | Implemented | lawful_intercept |
Interception is enforced in the dispatcher against ADMF-provisioned warrants — every message, every leg, every path. It is not opt-in from Python: the li.* script API is for visibility and for operator-driven SIPREC, and cannot prevent a warrant being actioned. |
| ETSI X1 provisioning (TS 103 221-1) | Implemented | lawful_intercept.x1 |
Conformant network element built to v1.23.1 with the TS 103 280 v2.19.1 dictionary; the ETSI XSDs ship in schemas/etsi/ and every message is validated against them in both directions at runtime. One application/xml endpoint (/X1/NE, configurable), mutual TLS with a mandatory client_ca, xsi:type dispatch over the X1Request/X1Response containers, per-message ErrorResponse with clause 6.7 codes, and admfIdentifier bound to the client certificate's CN (1030 on mismatch). Tasks key on the XID (a UUID, the same 16 bytes every X2/X3 PDU carries); target identifiers are the dictionary's (sipUri, telUri, e164Number, impu, impi, imsi, imei, IPv4/IPv6) and an unsupported one is refused 3010 by name rather than ignored. Destinations are modelled: a task delivers only to the DIDs it names, and a destination a task still references cannot be removed (7010). Validated end to end against sipgate's li-simulator-x1x2x3, an independent MIT implementation of the same specification, over real mutual TLS — provisioning, read-back, the 2010 and 7010 refusals, and teardown (scripts/run-tests.sh --li). |
| ETSI X1 network-element-to-ADMF direction | Implemented | lawful_intercept.x1.admf |
ReportNEIssue / ReportTaskIssue / ReportDestinationIssue, Keepalive on a timer, and GetAllDetails reconciliation at startup so a restart does not silently diverge the ADMF's view from the element's. Mutual TLS outbound. The inbound direction is validated against sipgate's simulator; the outbound reports are exercised by unit tests against a mock ADMF, not yet by a live one. |
| ETSI X2 IRI delivery (TS 103 221-2) | Implemented | lawful_intercept.x2 |
TS 103 221-2 clause 5 PDUs over TCP or TLS to the Mediation Function — the 40-octet mandatory header, conditional-attribute TLVs, and the SIP message carried verbatim as payload format 9. (Records used to be TS 102 232 PS-PDUs, which is the handover format an MDF emits onwards to the LEMF, not the interface into it.) Each record carries the task's XID, a non-zero per-session Correlation ID (a provisioned correlationID when there is one, else derived deterministically from the Call-ID so the media engine reaches the same value), and a direction measured against the target per clause 5.2.6. A task's records go to exactly the X2-capable DIDs it names, and to all of them. TLS is mutual and refuses rather than downgrading to plaintext. Backend-independent — works on every media.backend. Validated against a third-party dissector (scripts/validate_x2_pdu.sh) and end to end against sipgate's mediation function. |
| ETSI X3 CC delivery (TS 103 221-2) | Implemented | lawful_intercept.x3 |
Requires media.backend: siphon-rtp. Content framing lives in the media engine, so rtpengine and rtpproxy cannot deliver X3 at all: configuring lawful_intercept.x3 on either is refused at config load naming the backend, and an ActivateTask whose deliveryType is X3Only/X2andX3 is refused 3040 on such a node rather than accepted and silently delivering nothing. On the native backend siphon issues AttachX3 when a content warrant matches a dialog-forming request and DetachX3 at teardown, carrying the task's XID, the session's non-zero Correlation ID (the same value its X2 records carry, per clause 6) and the target leg — which is what TS 103 221-2 §5.2.6 measures each packet's direction against, so it is derived from which party the warrant matched rather than assumed. Engine X3Loss / an unclean X3Ended are raised to the ADMF as destination-level reports, which is what closes the loop: a mediation outage is reported, not merely survived. Needs siphon-rtp-proto >= 0.3.1. The attachment is made at the ACK rather than at the INVITE: interception is matched before the script runs, so at the INVITE the engine has not been offered the call yet and at the answer its second leg is not answered yet — the ACK is the first message with the session established on both legs. Validated end to end against sipgate's mediation function: a real call delivers X2 signalling and X3 content that decode as SIP and RTP respectively and share a correlation ID. |
| SIPREC recording (RFC 7866) | Implemented | lawful_intercept.siprec |
SIP Recording Server integration. A recording feature, not lawful interception: it produces no X2 record and is not tied to a provisioned warrant. |
Summary¶
| Category | Production | Implemented | Total |
|---|---|---|---|
| Transports | 4 (UDP, TCP, TLS, TLS 1.3) | 5 (WS, WSS, SCTP, mTLS, TLS 1.2) | 9 |
| Registrar | 7 (Redis, expires, max contacts, hooks, TTL slack, Service-Route, registrant) | 3 (memory, PG, Python, GRUU) | 10 |
| Authentication | 6 (HTTP/HA1, digest 401/407, anti-spoof, Diameter Cx, IMS AKA) | 3 (static, local Milenage AKA, SHA-256) | 9 |
| Security | 5 (rate limit, scanner, trusted CIDR, fail ban, APIBan) | 1 (IP ACLs) | 6 |
| NAT | 5 (rport, fix contact, fix register, script fixup, stale eviction) | 3 (keepalive, CRLF keepalive, flow tokens) | 8 |
| Media | 1 (RTPEngine NG) | 7 (LB, 4 profiles, custom profiles, v4↔v6 interworking) | 8 |
| Gateway routing | 3 (groups, round-robin, probes) | 4 (weighted, hash, failover, dynamic) | 7 |
| CDR | 0 | 5 (file, syslog, HTTP, register events, extra fields) | 5 |
| Tracing | 3 (HEP v3 UDP, agent ID, error suppression) | 2 (TCP, TLS) | 5 |
| Metrics | 8 (Prometheus, all gauges/counters/histograms) | 3 (admin health, stats, registrations) | 11 |
| Scripting | 14 (proxy, B2BUA, registrar, auth, gateway, cache, presence, logging, metrics, async, ...) | 3 (LI, timer, SDK) | 17 |
| 3GPP/IMS | 10 (Cx, Sh, Rx, peer mgmt, IMS AKA HSS-backed, IPsec, iFC, P/I/S-CSCF, Npcf) | 4 (Ro, Rf, local Milenage AKA, Nchf) | 14 |
| LI/Recording | 0 | 5 (X1, X2, X3, SIPREC, audit) | 5 |
| Control plane | 0 | 17 (handover, both connection modes, one-owner dispatch, deadline, resync, backpressure, call/media/header verbs, originate, REFER in + out, events, adapter API, SDKs, metrics) | 19 |
| Totals | ~66 | ~60 | ~129 |