Media & RTP profiles¶
SIPhon anchors and transforms media through a pluggable media engine — RTPEngine over its NG control protocol by default, or the native siphon-rtp engine (choosing and managing an engine). A profile is a named bundle of engine flags — SRTP↔RTP interworking, WebRTC, ICE handling, transcoding direction — that you select per call with one argument.
This page is the scripting recipe — the offer / answer / delete
lifecycle and the profile catalogue. It is identical for both backends. For
which engine to run and how to operate each one, see
Media engines: rtpengine vs siphon-rtp.
Config¶
# siphon.yaml
media:
rtpengine:
address: "127.0.0.1:22222" # NG control protocol (UDP)
timeout_ms: 1000
sdp_name: "SIPhon" # masks the endpoint identity in o=/s=
health_check_interval_secs: 5 # exported as siphon_rtpengine_instances_up
Multiple engines load-balance with weighted round-robin:
media:
rtpengine:
instances:
- { address: "10.0.0.1:22222", weight: 2 }
- { address: "10.0.0.2:22222", weight: 1 }
Choosing a media engine¶
SIPhon drives one of three media engines, chosen with media.backend:
media.backend |
Engine | Control transport |
|---|---|---|
rtpengine (default) |
RTPEngine | NG protocol, bencode over UDP |
siphon-rtp |
the in-house siphon-rtp engine | native JSON over a persistent TCP connection |
rtpproxy |
classic rtpproxy relay | text protocol over UDP |
Everything else on this page — the offer / answer / delete lifecycle, the
profiles, and the rtpengine scripting namespace — is identical for all
backends; only the transport underneath changes.
siphon-rtp is experimental
The siphon-rtp engine is pre-release, so this backend is experimental —
use the default rtpengine backend in production until it stabilises.
SIPREC/MPTY subscriptions are not yet implemented on siphon-rtp.
See Media engines: rtpengine vs siphon-rtp for the full
comparison, the media.siphon_rtp config, and how to run and operate each engine.
Classic rtpproxy (keep your existing relay)¶
Migrating an OpenSIPS / Kamailio / Sippy deployment? Point siphon at your existing
rtpproxy instead of standing up a new media engine — the script is unchanged, only
the config differs:
media:
backend: rtpproxy
rtpproxy:
address: "127.0.0.1:22222" # rtpproxy -s udp:<addr>
timeout_ms: 1000
retries: 2 # UDP retransmits (same cookie); rtpproxy de-dupes
# or several, for HA / weighted load-balancing (per-call-id affinity)
media:
backend: rtpproxy
rtpproxy:
instances:
- { address: "10.0.0.1:22222", weight: 2 }
- { address: "10.0.0.2:22222", weight: 1 }
siphon speaks rtpproxy's classic U/L/D protocol on the wire. Because rtpproxy
only hands back a relay port (it does not rewrite SDP itself), siphon rewrites the
c=/m= lines for you — per media stream, including held media (m=… 0). Profiles
still apply, but only the flags rtpproxy understands: a profile's
direction: ["internal","external"] becomes bridge mode (ie/ei) and an
asymmetric flag maps through; IPv6 is detected per stream. SRTP/DTLS/ICE flags are
ignored — rtpproxy is a plain RTP relay (use rtpengine or siphon-rtp for SRTP↔RTP,
WebRTC, or transcoding). A profile's
address_family is unsupported here too — rtpproxy's
6 modifier reports the family of the address the command carries, it cannot select
one for the relay — and siphon warns at boot naming any profile that sets it. It has
the same weighted round-robin + per-call-id affinity and per-instance V health
probes as the other backends.
rtpproxy is anchor-only
The extra rtpengine verbs — announcements / tones (play_media, play_dtmf),
gating (silence_media / block_media), DTMF events, and SIPREC/MPTY
subscriptions — are not available on the rtpproxy backend and raise a clear
error. They need rtpengine or siphon-rtp.
The offer / answer / delete lifecycle¶
Anchor the offer when the INVITE arrives, the answer when the 2xx comes back, and release on teardown. RTPEngine rewrites the SDP so media flows through it.
On a proxy:
from siphon import proxy, registrar, rtpengine
@proxy.on_request
async def route(request):
if request.in_dialog:
if request.method == "BYE":
await rtpengine.delete(request)
elif request.method == "INVITE" and request.body:
await rtpengine.offer(request, profile="srtp_to_rtp") # re-INVITE
request.loose_route() and request.relay()
return
contacts = registrar.lookup(request.ruri)
if request.method == "INVITE" and request.body:
await rtpengine.offer(request, profile="srtp_to_rtp")
request.record_route()
request.fork([c.uri for c in contacts])
@proxy.on_reply
async def reply_route(request, reply):
if 200 <= reply.status_code < 300 and reply.has_body("application/sdp"):
await rtpengine.answer(reply, profile="srtp_to_rtp")
reply.relay()
@proxy.on_cancel
async def cancel_route(request):
await rtpengine.delete(request) # release media for an abandoned call
On a B2BUA it's the same three calls in @b2bua.on_invite / on_answer /
on_bye (+ on_failure / on_cancel); pass call= to answer() so it reuses the
A-leg Call-ID that matched the offer (see the SBC recipe).
When the far side is not SIP¶
Sometimes the other side of a call is not a SIP agent: a media server your script
reaches over its own API, which terminates RTP itself and hands back an answer SDP.
There is no reply to pass to answer(), so pass the SDP. The coroutine resolves to
the rewritten SDP, which you send in your own 2xx:
@b2bua.on_invite
async def on_invite(call):
await rtpengine.offer(call, profile="rtp_passthrough") # rewrites call.body
far_sdp = await my_media_server.connect(call.body) # your transport, not SIP
sdp = await rtpengine.answer(call, sdp=far_sdp)
call.answer(200, "OK", body=sdp, content_type="application/sdp")
call names the offer being answered. From a handler that only has the
identifiers, pass (call_id, from_tag) instead; a bare call_id is refused because
it names no from-tag. to_tag= names the answering party to the engine. Leave it
out and siphon reuses the tag an earlier answer on the call recorded, so a later
re-offer and re-answer reach the same party. The offer's profile still decides the
media work, transcoding included when the two sides share no codec. No source
address is carried for the far side (siphon never heard from it), so a profile's
received_from does not gate its media.
Always release
offer without a matching delete leaks an RTPEngine session until its
inactivity timeout. Handle every teardown path — on_bye, on_failure,
on_cancel (proxy: @proxy.on_cancel) — or media lingers.
A failed answer in @b2bua.on_answer fails the call
If rtpengine.answer() raises there, siphon does not connect the call: the
answered B-leg is ACKed and BYEd, on_failure fires with 500, and no
answer-time charging is reported. Unless the failure handler routes the call
somewhere else, the caller gets that 500. That is deliberate — the
B-leg has answered but the A-leg has not yet, so this is the last point at
which a call with no media path can still be stopped rather than billed.
Catch the exception yourself only if you can actually recover; swallowing it
to keep the call up gives you a connected call that carries no audio in
either direction. call.terminate() from the same handler does the same
thing explicitly.
Built-in profiles¶
| Profile | Interworking |
|---|---|
rtp_passthrough |
Plain RTP both sides — anchoring only (the default) |
srtp_to_rtp |
SRTP UE ↔ RTP core (VoLTE/secure access ↔ trunk) |
rtp_to_srtp |
The reverse pairing — RTP access ↔ SRTP core |
ws_to_rtp |
WebSocket UE (RTP/AVPF + ICE) ↔ RTP core |
wss_to_rtp |
Secure WebSocket (DTLS-SRTP/AVPF + ICE) ↔ RTP core |
srs_recording |
Recording sink — plain RTP, media handover + port latching |
siprec_src |
SIPREC SRC subscription leg toward the recorder |
voice_ai |
Plain RTP toward the caller, audio bridged to a WebSocket AI backend |
ws_to_rtp / wss_to_rtp are what make a WebRTC gateway work — terminate the
browser's DTLS-SRTP + ICE on one side, plain RTP toward your core on the other.
Custom profiles¶
Define your own under media.profiles — any RTPEngine flag, per direction:
media:
profiles:
srtp_to_srtp:
offer:
transport_protocol: "RTP/SAVP"
ice: "remove"
replace: ["origin"]
direction: ["external", "internal"]
answer:
transport_protocol: "RTP/SAVP"
ice: "remove"
replace: ["origin"]
direction: ["internal", "external"]
IPv4 and IPv6 interworking¶
address_family pins the family the engine allocates its own relay endpoints
in for that side of the call. Leave it unset (the default) and the engine follows
the offered SDP, which gives you a single-family relay — fine until one side is
v6-only. Set it per direction to bridge, e.g. a v6 VoLTE access leg reaching a v4
core:
media:
profiles:
v6_access_to_v4_core:
offer: # toward the core: hand it a v4 endpoint
replace: ["origin"]
address_family: "IP4"
answer: # back toward the v6 UE
replace: ["origin"]
address_family: "IP6"
The value is the SDP addrtype spelling, IP4 or IP6 (ipv4 / ipv6 are
accepted and normalised; anything else fails the config load, because a media
engine ignores an unknown family silently and you would get a relay in the wrong
family with no error). The engine needs an interface configured in the target
family — rtpengine's interface= must list both, otherwise it has nothing to
allocate from.
Works on rtpengine (sent as the dedicated address family NG key) and
siphon-rtp (the address_family control field). The classic rtpproxy
backend has no equivalent and logs a warning at boot if a profile sets it.
Bridge a leg's audio to a WebSocket server¶
ws_uri hands a leg's audio to an external WebSocket media server instead of a
far SIP leg: the engine dials the URI and relays the leg's RTP to it as L16
(decode → uplink, downlink → encode). The WS server is that leg's far side, so
this is the shape a call answered by a speech backend takes — pair it with
rtpengine.answer_local, which
synthesises the 2xx answer with the engine as the far side.
media:
backend: siphon-rtp # required — see the capability table below
siphon_rtp:
address: "127.0.0.1:9000"
profiles:
voice_ai: # overrides the built-in of the same name
offer: &voice_ai_flags
transport_protocol: "RTP/AVP"
ice: "remove"
dtls: "off"
replace: ["origin"]
ws_uri: "wss://ai.example.com/stream/{call_id}"
ws_vad: true # emit speech_started / speech_stopped edges
ws_barge_in: true # cut playout locally on the caller's speech
ws_vad_engine: neural # "is this speech", not "is this loud"
ws_vad_min_speech_ms: 100 # leading run before the speech-start edge
ws_vad_threshold: 2000000
ws_vad_hangover_ms: 300
noise_suppression: true
echo_cancellation: true
echo_delay_search_ms: 400
received_from: true # gate on the real post-NAT source
answer: *voice_ai_flags
The URI supports {call_id}, {from_tag}, {from_user} and {to_user},
expanded per call. An unrecognised placeholder fails rather than passing through
as a literal, so a typo cannot reach the engine as part of the URI path.
When the endpoint depends on something only the script knows — a session token,
a tenant lookup — pass it per call instead. It wins over the profile's own value,
and is recorded on the media session so a later answer reuses the same bridge
without repeating it:
@b2bua.on_invite
async def on_invite(call):
sdp = await rtpengine.answer_local(
call,
profile="voice_ai",
ws_uri=f"wss://ai.example.com/stream?token={await mint_token(call.call_id)}",
)
if sdp is not None:
call.answer(200, "OK", body=sdp, content_type="application/sdp")
The built-in voice_ai profile sets the DSP and VAD flags (including
ws_vad_engine: neural, ws_vad_min_speech_ms: 100 and received_from) but
deliberately leaves ws_uri unset — there is no sensible default endpoint, so
supply it in YAML or per call as above. An entry of the same name under
media.profiles replaces the built-in rather than merging with it, so
restate every flag the override still wants.
ws_uri, the ws_* knobs, noise_suppression and the echo_* knobs are
siphon-rtp only. siphon refuses to start if a media.profiles entry sets
one on another backend, and a script naming such a profile gets a ValueError
naming the field — see media engines for the full
capability table.
Gate media ingress to the signalling source¶
received_from: true carries the real post-NAT source IP siphon saw the request
arrive from, and the engine gates that leg's ingress to it. For a NATed UA whose
c= line advertises an unroutable private address, that is a tighter
RTPBleed source gate than the signalled address could give. Only the IP is
carried — media and signalling ports differ, so the port is never gated.
media:
profiles:
nated_access:
offer:
replace: ["origin"]
received_from: true
answer:
replace: ["origin"]
received_from: true
Each party is pinned to its own signalling source, by its own half of the
profile: the offer half is the caller's and the answer half the callee's,
whichever of them sends the SDP. rtpengine.offer(request) carries the
request's source, rtpengine.answer(reply) the source the reply arrived from
(never the caller's, with or without call=), and a re-INVITE or UPDATE from
the callee is pinned by the answer half although it travels as an offer. The
same holds on a delayed offer, where the callee offers in its 2xx and the caller
answers in its ACK. Set it on one half only when just one side is behind NAT.
The rest of a half works the other way round, and it helps to keep the two
apart. received_from is about the party whose SDP a command carries.
Everything that shapes SDP (transport_protocol, direction, ICE, DTLS, codec
handling) is about the party the rewritten SDP is sent to: the offer half
is what the callee is sent and the answer half what the caller is sent. That
too holds for the life of the call. When the callee re-offers (a hold, say),
its offer is relayed to the caller under the answer half and the caller's
answer goes back to the callee under the offer half, so with srtp_to_rtp
the SRTP side is offered SRTP and the plain side plain RTP whoever sends the
re-INVITE. A script keeps passing the same profile= to rtpengine.offer() and
rtpengine.answer() on every request and reply, as in the examples above.
Off by default, because it is wrong for a deployment whose media legitimately
arrives from a different address than its signalling (a separate media gateway,
or a carrier that splits the two). Honoured by rtpengine (the received from
NG key) and siphon-rtp; rtpproxy has no equivalent and fails the config
load.
rtcp_mux takes the same RFC 5761 directives rtpengine does — offer,
require, demux, accept, reject, remove — to override the mux decision
the engine would derive from the offered SDP. Empty (the default) mirrors the
offer. An unknown token fails the config load rather than being silently dropped
by the engine.
Shape the SDP yourself¶
For codec filtering, hold, or attribute tweaks without RTPEngine, use the sdp
namespace:
from siphon import sdp
s = sdp.parse(request)
for m in s.media:
if m.media_type == "audio":
s.filter_codecs(["PCMU", "PCMA"]) # keep only G.711
# m.port = 0 # ... or put audio on hold
s.apply(request)
More media control¶
The rtpengine namespace also drives announcements and tones (play_media,
play_dtmf), gating (silence_media / block_media), DTMF events (@rtpengine.on_dtmf),
and conference/MPTY subscriptions — useful for IVR, MMTel announcements, and recording.
React to a dead media path¶
The media engine reaps a call whose media stops flowing (no packets past its
inactivity window). Handle @rtpengine.on_media_timeout to release the per-call
state that no BYE will now clear — Rx/N5 QoS sessions, offline charging, dialog
or session-store entries. It is the media-path analogue of the abandoned-call
teardown @proxy.on_cancel / @b2bua.on_cancel cover.
from siphon import rtpengine, log
@rtpengine.on_media_timeout
async def media_gone(call_id, from_tag):
log.warn(f"media timeout on {call_id} — releasing call state")
# The release calls are awaitable, so the handler is `async def`:
# e.g. await diameter.rx_str(session_id) / await sbi.delete_session(...)
# / cdr.write(...)
Filter to a specific call with @rtpengine.on_media_timeout(call_id=..., from_tag=...),
the same shape as @rtpengine.on_dtmf.
siphon-rtp only, for now
This event is delivered by the native siphon-rtp backend, which pushes it
over its control connection. The rtpengine backend's event log carries only
DTMF, so @rtpengine.on_media_timeout does not fire under rtpengine yet —
see Media engines.
See also¶
- Real examples:
examples/proxy_rtpengine.py,examples/b2bua_rtpengine.py. - SBC (B2BUA) — media anchoring in a topology-hiding SBC.