Skip to content

Media & RTP profiles

SIPhon anchors and transforms media through a pluggable media engine — RTPEngine over its NG control protocol by default, or the native siphon-rtp engine (choosing and managing an engine). A profile is a named bundle of engine flags — SRTP↔RTP interworking, WebRTC, ICE handling, transcoding direction — that you select per call with one argument.

This page is the scripting recipe — the offer / answer / delete lifecycle and the profile catalogue. It is identical for both backends. For which engine to run and how to operate each one, see Media engines: rtpengine vs siphon-rtp.

Config

# siphon.yaml
media:
  rtpengine:
    address: "127.0.0.1:22222"     # NG control protocol (UDP)
    timeout_ms: 1000
  sdp_name: "SIPhon"               # masks the endpoint identity in o=/s=
  health_check_interval_secs: 5    # exported as siphon_rtpengine_instances_up

Multiple engines load-balance with weighted round-robin:

media:
  rtpengine:
    instances:
      - { address: "10.0.0.1:22222", weight: 2 }
      - { address: "10.0.0.2:22222", weight: 1 }

Choosing a media engine

SIPhon drives one of three media engines, chosen with media.backend:

media.backend Engine Control transport
rtpengine (default) RTPEngine NG protocol, bencode over UDP
siphon-rtp the in-house siphon-rtp engine native JSON over a persistent TCP connection
rtpproxy classic rtpproxy relay text protocol over UDP

Everything else on this page — the offer / answer / delete lifecycle, the profiles, and the rtpengine scripting namespace — is identical for all backends; only the transport underneath changes.

siphon-rtp is experimental

The siphon-rtp engine is pre-release, so this backend is experimental — use the default rtpengine backend in production until it stabilises. SIPREC/MPTY subscriptions are not yet implemented on siphon-rtp.

See Media engines: rtpengine vs siphon-rtp for the full comparison, the media.siphon_rtp config, and how to run and operate each engine.

Classic rtpproxy (keep your existing relay)

Migrating an OpenSIPS / Kamailio / Sippy deployment? Point siphon at your existing rtpproxy instead of standing up a new media engine — the script is unchanged, only the config differs:

media:
  backend: rtpproxy
  rtpproxy:
    address: "127.0.0.1:22222"     # rtpproxy -s udp:<addr>
    timeout_ms: 1000
    retries: 2                     # UDP retransmits (same cookie); rtpproxy de-dupes

# or several, for HA / weighted load-balancing (per-call-id affinity)
media:
  backend: rtpproxy
  rtpproxy:
    instances:
      - { address: "10.0.0.1:22222", weight: 2 }
      - { address: "10.0.0.2:22222", weight: 1 }

siphon speaks rtpproxy's classic U/L/D protocol on the wire. Because rtpproxy only hands back a relay port (it does not rewrite SDP itself), siphon rewrites the c=/m= lines for you — per media stream, including held media (m=… 0). Profiles still apply, but only the flags rtpproxy understands: a profile's direction: ["internal","external"] becomes bridge mode (ie/ei) and an asymmetric flag maps through; IPv6 is detected per stream. SRTP/DTLS/ICE flags are ignored — rtpproxy is a plain RTP relay (use rtpengine or siphon-rtp for SRTP↔RTP, WebRTC, or transcoding). A profile's address_family is unsupported here too — rtpproxy's 6 modifier reports the family of the address the command carries, it cannot select one for the relay — and siphon warns at boot naming any profile that sets it. It has the same weighted round-robin + per-call-id affinity and per-instance V health probes as the other backends.

rtpproxy is anchor-only

The extra rtpengine verbs — announcements / tones (play_media, play_dtmf), gating (silence_media / block_media), DTMF events, and SIPREC/MPTY subscriptions — are not available on the rtpproxy backend and raise a clear error. They need rtpengine or siphon-rtp.

The offer / answer / delete lifecycle

Anchor the offer when the INVITE arrives, the answer when the 2xx comes back, and release on teardown. RTPEngine rewrites the SDP so media flows through it.

On a proxy:

from siphon import proxy, registrar, rtpengine

@proxy.on_request
async def route(request):
    if request.in_dialog:
        if request.method == "BYE":
            await rtpengine.delete(request)
        elif request.method == "INVITE" and request.body:
            await rtpengine.offer(request, profile="srtp_to_rtp")  # re-INVITE
        request.loose_route() and request.relay()
        return

    contacts = registrar.lookup(request.ruri)
    if request.method == "INVITE" and request.body:
        await rtpengine.offer(request, profile="srtp_to_rtp")
    request.record_route()
    request.fork([c.uri for c in contacts])

@proxy.on_reply
async def reply_route(request, reply):
    if 200 <= reply.status_code < 300 and reply.has_body("application/sdp"):
        await rtpengine.answer(reply, profile="srtp_to_rtp")
    reply.relay()

@proxy.on_cancel
async def cancel_route(request):
    await rtpengine.delete(request)   # release media for an abandoned call

On a B2BUA it's the same three calls in @b2bua.on_invite / on_answer / on_bye (+ on_failure / on_cancel); pass call= to answer() so it reuses the A-leg Call-ID that matched the offer (see the SBC recipe).

When the far side is not SIP

Sometimes the other side of a call is not a SIP agent: a media server your script reaches over its own API, which terminates RTP itself and hands back an answer SDP. There is no reply to pass to answer(), so pass the SDP. The coroutine resolves to the rewritten SDP, which you send in your own 2xx:

@b2bua.on_invite
async def on_invite(call):
    await rtpengine.offer(call, profile="rtp_passthrough")   # rewrites call.body
    far_sdp = await my_media_server.connect(call.body)       # your transport, not SIP
    sdp = await rtpengine.answer(call, sdp=far_sdp)
    call.answer(200, "OK", body=sdp, content_type="application/sdp")

call names the offer being answered. From a handler that only has the identifiers, pass (call_id, from_tag) instead; a bare call_id is refused because it names no from-tag. to_tag= names the answering party to the engine. Leave it out and siphon reuses the tag an earlier answer on the call recorded, so a later re-offer and re-answer reach the same party. The offer's profile still decides the media work, transcoding included when the two sides share no codec. No source address is carried for the far side (siphon never heard from it), so a profile's received_from does not gate its media.

Always release

offer without a matching delete leaks an RTPEngine session until its inactivity timeout. Handle every teardown path — on_bye, on_failure, on_cancel (proxy: @proxy.on_cancel) — or media lingers.

A failed answer in @b2bua.on_answer fails the call

If rtpengine.answer() raises there, siphon does not connect the call: the answered B-leg is ACKed and BYEd, on_failure fires with 500, and no answer-time charging is reported. Unless the failure handler routes the call somewhere else, the caller gets that 500. That is deliberate — the B-leg has answered but the A-leg has not yet, so this is the last point at which a call with no media path can still be stopped rather than billed. Catch the exception yourself only if you can actually recover; swallowing it to keep the call up gives you a connected call that carries no audio in either direction. call.terminate() from the same handler does the same thing explicitly.

Built-in profiles

Profile Interworking
rtp_passthrough Plain RTP both sides — anchoring only (the default)
srtp_to_rtp SRTP UE ↔ RTP core (VoLTE/secure access ↔ trunk)
rtp_to_srtp The reverse pairing — RTP access ↔ SRTP core
ws_to_rtp WebSocket UE (RTP/AVPF + ICE) ↔ RTP core
wss_to_rtp Secure WebSocket (DTLS-SRTP/AVPF + ICE) ↔ RTP core
srs_recording Recording sink — plain RTP, media handover + port latching
siprec_src SIPREC SRC subscription leg toward the recorder
voice_ai Plain RTP toward the caller, audio bridged to a WebSocket AI backend

ws_to_rtp / wss_to_rtp are what make a WebRTC gateway work — terminate the browser's DTLS-SRTP + ICE on one side, plain RTP toward your core on the other.

Custom profiles

Define your own under media.profiles — any RTPEngine flag, per direction:

media:
  profiles:
    srtp_to_srtp:
      offer:
        transport_protocol: "RTP/SAVP"
        ice: "remove"
        replace: ["origin"]
        direction: ["external", "internal"]
      answer:
        transport_protocol: "RTP/SAVP"
        ice: "remove"
        replace: ["origin"]
        direction: ["internal", "external"]
await rtpengine.offer(request, profile="srtp_to_srtp")

IPv4 and IPv6 interworking

address_family pins the family the engine allocates its own relay endpoints in for that side of the call. Leave it unset (the default) and the engine follows the offered SDP, which gives you a single-family relay — fine until one side is v6-only. Set it per direction to bridge, e.g. a v6 VoLTE access leg reaching a v4 core:

media:
  profiles:
    v6_access_to_v4_core:
      offer:                     # toward the core: hand it a v4 endpoint
        replace: ["origin"]
        address_family: "IP4"
      answer:                    # back toward the v6 UE
        replace: ["origin"]
        address_family: "IP6"

The value is the SDP addrtype spelling, IP4 or IP6 (ipv4 / ipv6 are accepted and normalised; anything else fails the config load, because a media engine ignores an unknown family silently and you would get a relay in the wrong family with no error). The engine needs an interface configured in the target family — rtpengine's interface= must list both, otherwise it has nothing to allocate from.

Works on rtpengine (sent as the dedicated address family NG key) and siphon-rtp (the address_family control field). The classic rtpproxy backend has no equivalent and logs a warning at boot if a profile sets it.

Bridge a leg's audio to a WebSocket server

ws_uri hands a leg's audio to an external WebSocket media server instead of a far SIP leg: the engine dials the URI and relays the leg's RTP to it as L16 (decode → uplink, downlink → encode). The WS server is that leg's far side, so this is the shape a call answered by a speech backend takes — pair it with rtpengine.answer_local, which synthesises the 2xx answer with the engine as the far side.

media:
  backend: siphon-rtp          # required — see the capability table below
  siphon_rtp:
    address: "127.0.0.1:9000"
  profiles:
    voice_ai:                  # overrides the built-in of the same name
      offer: &voice_ai_flags
        transport_protocol: "RTP/AVP"
        ice: "remove"
        dtls: "off"
        replace: ["origin"]
        ws_uri: "wss://ai.example.com/stream/{call_id}"
        ws_vad: true           # emit speech_started / speech_stopped edges
        ws_barge_in: true      # cut playout locally on the caller's speech
        ws_vad_engine: neural  # "is this speech", not "is this loud"
        ws_vad_min_speech_ms: 100  # leading run before the speech-start edge
        ws_vad_threshold: 2000000
        ws_vad_hangover_ms: 300
        noise_suppression: true
        echo_cancellation: true
        echo_delay_search_ms: 400
        received_from: true    # gate on the real post-NAT source
      answer: *voice_ai_flags

The URI supports {call_id}, {from_tag}, {from_user} and {to_user}, expanded per call. An unrecognised placeholder fails rather than passing through as a literal, so a typo cannot reach the engine as part of the URI path.

When the endpoint depends on something only the script knows — a session token, a tenant lookup — pass it per call instead. It wins over the profile's own value, and is recorded on the media session so a later answer reuses the same bridge without repeating it:

@b2bua.on_invite
async def on_invite(call):
    sdp = await rtpengine.answer_local(
        call,
        profile="voice_ai",
        ws_uri=f"wss://ai.example.com/stream?token={await mint_token(call.call_id)}",
    )
    if sdp is not None:
        call.answer(200, "OK", body=sdp, content_type="application/sdp")

The built-in voice_ai profile sets the DSP and VAD flags (including ws_vad_engine: neural, ws_vad_min_speech_ms: 100 and received_from) but deliberately leaves ws_uri unset — there is no sensible default endpoint, so supply it in YAML or per call as above. An entry of the same name under media.profiles replaces the built-in rather than merging with it, so restate every flag the override still wants.

ws_uri, the ws_* knobs, noise_suppression and the echo_* knobs are siphon-rtp only. siphon refuses to start if a media.profiles entry sets one on another backend, and a script naming such a profile gets a ValueError naming the field — see media engines for the full capability table.

Gate media ingress to the signalling source

received_from: true carries the real post-NAT source IP siphon saw the request arrive from, and the engine gates that leg's ingress to it. For a NATed UA whose c= line advertises an unroutable private address, that is a tighter RTPBleed source gate than the signalled address could give. Only the IP is carried — media and signalling ports differ, so the port is never gated.

media:
  profiles:
    nated_access:
      offer:
        replace: ["origin"]
        received_from: true
      answer:
        replace: ["origin"]
        received_from: true

Each party is pinned to its own signalling source, by its own half of the profile: the offer half is the caller's and the answer half the callee's, whichever of them sends the SDP. rtpengine.offer(request) carries the request's source, rtpengine.answer(reply) the source the reply arrived from (never the caller's, with or without call=), and a re-INVITE or UPDATE from the callee is pinned by the answer half although it travels as an offer. The same holds on a delayed offer, where the callee offers in its 2xx and the caller answers in its ACK. Set it on one half only when just one side is behind NAT.

The rest of a half works the other way round, and it helps to keep the two apart. received_from is about the party whose SDP a command carries. Everything that shapes SDP (transport_protocol, direction, ICE, DTLS, codec handling) is about the party the rewritten SDP is sent to: the offer half is what the callee is sent and the answer half what the caller is sent. That too holds for the life of the call. When the callee re-offers (a hold, say), its offer is relayed to the caller under the answer half and the caller's answer goes back to the callee under the offer half, so with srtp_to_rtp the SRTP side is offered SRTP and the plain side plain RTP whoever sends the re-INVITE. A script keeps passing the same profile= to rtpengine.offer() and rtpengine.answer() on every request and reply, as in the examples above.

Off by default, because it is wrong for a deployment whose media legitimately arrives from a different address than its signalling (a separate media gateway, or a carrier that splits the two). Honoured by rtpengine (the received from NG key) and siphon-rtp; rtpproxy has no equivalent and fails the config load.

rtcp_mux takes the same RFC 5761 directives rtpengine does — offer, require, demux, accept, reject, remove — to override the mux decision the engine would derive from the offered SDP. Empty (the default) mirrors the offer. An unknown token fails the config load rather than being silently dropped by the engine.

Shape the SDP yourself

For codec filtering, hold, or attribute tweaks without RTPEngine, use the sdp namespace:

from siphon import sdp

s = sdp.parse(request)
for m in s.media:
    if m.media_type == "audio":
        s.filter_codecs(["PCMU", "PCMA"])   # keep only G.711
        # m.port = 0                          # ... or put audio on hold
s.apply(request)

More media control

The rtpengine namespace also drives announcements and tones (play_media, play_dtmf), gating (silence_media / block_media), DTMF events (@rtpengine.on_dtmf), and conference/MPTY subscriptions — useful for IVR, MMTel announcements, and recording.

React to a dead media path

The media engine reaps a call whose media stops flowing (no packets past its inactivity window). Handle @rtpengine.on_media_timeout to release the per-call state that no BYE will now clear — Rx/N5 QoS sessions, offline charging, dialog or session-store entries. It is the media-path analogue of the abandoned-call teardown @proxy.on_cancel / @b2bua.on_cancel cover.

from siphon import rtpengine, log

@rtpengine.on_media_timeout
async def media_gone(call_id, from_tag):
    log.warn(f"media timeout on {call_id} — releasing call state")
    # The release calls are awaitable, so the handler is `async def`:
    # e.g. await diameter.rx_str(session_id) / await sbi.delete_session(...)
    # / cdr.write(...)

Filter to a specific call with @rtpengine.on_media_timeout(call_id=..., from_tag=...), the same shape as @rtpengine.on_dtmf.

siphon-rtp only, for now

This event is delivered by the native siphon-rtp backend, which pushes it over its control connection. The rtpengine backend's event log carries only DTMF, so @rtpengine.on_media_timeout does not fire under rtpengine yet — see Media engines.

See also