Hardening & security¶
A SIP port on the public internet gets scanned within minutes. This recipe collects the layers SIPhon gives you — most are config, a few are one-liners in a script.
1. Drop abuse before it costs you (config)¶
The security: block runs before any SIP parsing or scripting, so banned/garbage
traffic never reaches your handlers:
security:
rate_limit:
window_secs: 10
max_requests: 30 # per source IP per window
ban_duration_secs: 3600
scanner_block:
user_agents: ["sipvicious", "friendly-scanner", "VaxSip", "sipcli"]
trusted_cidrs: ["10.0.0.0/8"] # own infra: never rate-limited, never banned,
# never refused by connection_limits
connection_limits: # always on — every field defaults
max_handshakes_per_source: 32
max_handshakes: 1024
max_connections_per_source: 256
max_connections: 16384
failed_auth_ban: # auto-ban at accept (UDP/TCP/TLS/WS/SCTP)
threshold: 10 # weighted failures in window_secs → ban
window_secs: 600
ban_duration_secs: 3600 # expiry slides on continued abuse
max_ban_duration_secs: 86400 # cap on the slide (default 24 ×
# ban_duration_secs)
strong_signal_weight: 3 # weight of a high-confidence abuse signal
missing_credentials_weight: 0 # default: a credential-less request is the
# RFC-mandated first leg, not evidence
apiban: # optional: APIBAN community blocklist
api_key: "your-api-key"
interval_secs: 300
ban_ttl_secs: 604800 # 7 days, matching the feed's own release
# policy. 0 = never expire.
trusted_cidrs covers the feed too: an address listed by APIBAN that matches a
trusted CIDR is dropped as the feed is ingested, so it reaches neither the
userspace ACL nor the kernel set. Put your own trunks, monitoring and management
addresses there — a community blocklist has no way to know they're yours, and
the kernel drop is port-agnostic, so a listed management address would cost you
ssh along with the trunk.
How the scoring works¶
failed_auth_ban is a confidence-weighted counter, not a flat fail2ban tally.
Every abuse signal from a source IP adds to a per-IP score within window_secs;
crossing threshold bans the IP for ban_duration_secs. Signals are weighted by
how hard they are to fake:
| Signal | Score |
|---|---|
| INVITE server-transaction timeout (never ACKed) | 1 |
| Failed or timed-out TLS handshake | 1 |
| Wrong password, a username the auth backend denied, or a forged/stale/replayed digest nonce | strong_signal_weight (default 3) |
| Non-SIP bytes on a TCP/TLS stream | strong_signal_weight |
| Rejected WebSocket upgrade on a WS/WSS port | strong_signal_weight |
Scanner User-Agent (scanner_block) |
strong_signal_weight |
| 401/407 challenge because the request carried no credentials | missing_credentials_weight (default 0 — not counted) |
| A credential check the auth backend could not answer | never counted |
Signals carrying present-but-wrong credentials, or garbage over TCP, score high because they are unambiguous: the source IP is validated by the three-way handshake so it cannot be spoofed, and a legitimate client never trips them. A successful authentication resets the score to zero, so a subscriber who mistypes a password twice then logs in is never banned, while an IP spraying garbage is banned 3× faster than one just rattling doorknobs.
The two handshake rows are not the same signal, and the split is deliberate. A
failed TLS handshake can come from a benign peer — a client that doesn't trust your
chain, an old cipher suite, an L4 probe — so it scores 1, and a certificate rollover
doesn't become a ban wave. A rejected WebSocket upgrade is a different animal: the
peer already completed TLS on a SIP-over-WebSocket port and then sent something that
isn't an upgrade at all. RFC 7118 §5 leaves a conforming client no way to do that,
so it scores as strong and bans in a couple of probes. The practical consequence: an
external HTTP uptime monitor pointed at a WS/WSS port will ban itself. Put its source
in trusted_cidrs.
A ban that slides¶
An active ban's expiry is not fixed. Every further abuse signal from an
already-banned source pushes it out to a full ban_duration_secs from that
signal. Without this, a scanner that trips the threshold and then keeps hammering
gets the rest of its run for free and walks out on the original schedule no matter
how hard it leaned on the box in between.
max_ban_duration_secs caps how far the slide can go, measured from the moment the
ban was raised, and it is the part that matters. Uncapped, a source stuck in a retry
loop is banned forever — and behind CGNAT that address speaks for every other
subscriber on the NAT, none of whom did anything. The default of 24 ×
ban_duration_secs holds a real scanner across a working day while letting a wrong
verdict age out on its own. Set it equal to ban_duration_secs for a fixed TTL.
The rate-limit ban (rate_limit) deliberately does not slide. Being over a rate
limit is a capacity verdict, not evidence of intent, so a client that keeps retrying
through its ban still serves it out on schedule.
The last two rows of the table are the ones worth understanding, because both were once counted and both banned real subscribers:
- A request with no credentials is not evidence. RFC 3261 §22.2 makes it the
opening leg of challenge-response — every client sends one before it has a nonce.
Counting it means a handset stuck in a retry loop earns an hour-long ban, and
behind CGNAT that address is shared, so the ban lands on every subscriber behind
it. Volume still shows in
siphon_auth_failures_total; setmissing_credentials_weight: 1if you want the old scoring back. - An auth-backend outage is not an attack. A
GETto your credential endpoint that times out tells you nothing about the peer. It used to be indistinguishable from a wrong password, so two REGISTER retries during an outage banned the subscriber — exactly when every subscriber is retrying. Alert onsiphon_auth_backend_errors_totalinstead; a non-zero rate means authentication is failing into 401s for everyone.
Bounding what one source can spend¶
connection_limits is a separate, always-on layer, and it covers what the ban
counter structurally cannot: a source that opens 50 TLS connections at once and
completes none of them never produces a completed failure to count, while each
connection burns a real handshake and pins a task for the full 10 s handshake
timeout.
Two ceilings, because the resources differ. An in-flight handshake is CPU held
briefly and no legitimate client has many at once, so that one is tight (32 per
source). An established connection is a socket held until the peer leaves or the
300 s idle timeout reaps it, and a busy NAT legitimately holds many, so that one is
loose (256 per source). Each has a global twin for distributed floods. 0 disables
a ceiling; trusted_cidrs are exempt from all of them.
Refused connections are dropped silently and not banned — hitting a concurrency ceiling is a capacity fact, not proof of intent, and a NAT whose UEs all re-register after a network flap looks exactly like a flood.
Carrier NAT and max_connections_per_source
A CGNAT pool or a large enterprise NAT can legitimately front more registrations
from one address than the 256 default allows, and every one of them is a paying
subscriber. The default is a runaway detector, not a policy — raise it, or set
0, wherever that is your topology. Watch
siphon_connections_refused_total{reason="connections_per_source"}: it tells you
the ceiling is binding on real traffic before the support tickets do.
siphon_stream_connections_active and siphon_handshakes_in_flight are what you
size against.
Bans are enforced at recv()/accept() — before any SIP parsing — and expire on
their own. trusted_cidrs are exempt from scoring entirely, so put your load
balancers and health checks there.
They are also re-checked once a TLS or WebSocket handshake completes, which closes a
window the accept-time check cannot: a scanner opens a burst of connections at once,
one of them trips the threshold, and every sibling already past accept() would
otherwise be served to completion because nothing looks at the ban again. The
re-check drops the rest of the burst with the connection that earned the ban.
And a new ban closes the connections its source already holds open — every live
TCP, TLS, WS, WSS and SCTP connection from that address, logged one line each. Both
checks above run once per connection, so a client that never reconnects used to keep
the connection it had and go on guessing passwords on it for the whole ban, while
every other client behind the same address was refused at accept(). A banned source
now has to come back through accept(), where it is refused until the ban expires.
Drop bans in the kernel
With security.firewall, every ban is also pushed to a
kernel nf_tables set, so abusive sources are dropped before they reach
SIPhon — real defense against volume, not just userspace politeness.
In a script, you can also rate-limit a specific flow:
if not proxy.rate_limit(request, window_secs=1, max_requests=5):
return # silently drop — don't fingerprint the server
Letting a front speak for the client¶
Everything above keys on the source IP, so everything above quietly stops working
when a front terminates the connection — an L7 or TLS-terminating proxy,
HAProxy in tcp mode with its own certificate, an Ingress controller. The front
opens its own connection, so that is the only address SIPhon sees: the ban store
bans the front or nobody, trusted_cidrs either exempts every client behind it
or none of them, and from_gateway() stops telling you which side a call came
from. proxy_protocol on the listener restores the client address by reading it
out of the header the front sends:
listen:
tls:
- address: "198.51.100.10:5061"
proxy_protocol:
from: ["198.51.100.7/32"] # the front, and nothing else
from is a security control, not a convenience. A PROXY header lets its
sender claim to be any address on the internet — including one of your
trusted_cidrs, which is exactly how an unrestricted PROXY listener becomes a
one-line bypass of every layer on this page. So from is mandatory, has no
default, and a listener with an empty or unparseable list is refused at config
load rather than started permissive. Keep it as narrow as the front really is: a
/32 per front, not the subnet it lives in.
It deliberately does not inherit trusted_cidrs. That is the obvious
shortcut and it is the wrong one. trusted_cidrs means "exempt from abuse
controls" to all four of its consumers, and it is where this page has just told
you to put monitoring boxes, health-check probes and trunks. Inheriting it would
hand source-address forgery rights to every one of them, on the strength of a
decision made for an entirely different reason. The two lists answer different
questions and stay separate.
Two failure directions are closed on purpose, because both would otherwise be silent:
- A connection from outside
from, and a connection on an enabled listener that opens with anything other than a PROXY header, are dropped — never attributed to the front. Neither credits the ban store: a second front nobody added to the list is likelier than an attack, and banning your own ingress is a worse outage than the misconfiguration. - A PROXY header arriving on a listener with the option off is recognised and
refused with a log naming the listener. It used to score as non-SIP bytes, i.e.
strong_signal_weight— SIPhon banning its own load balancer within a few connections.
The header is cleartext and arrives ahead of the TLS ClientHello, so it is read before the handshake. That is what lets the front terminate the subscriber's TLS and open its own to SIPhon with no hop carrying cleartext SIP. Stream listeners only; the full reference is in Transports.
2. Drop malformed traffic (script)¶
proxy.sanity_check() runs the RFC 4475 semantic checks (mandatory headers, CSeq,
Content-Length). Drop failures silently so scanners learn nothing:
@proxy.on_request
def route(request):
if not request.in_dialog and not proxy.sanity_check(request):
return # silent drop
...
Silent drop is intentional
Returning from a handler without reply()/relay()/reject() sends no response.
For rate-limit and scanner blocking that's the point — a 403 would confirm the
server exists. Don't "helpfully" reply.
3. Encrypt the signalling (config)¶
listen:
tls: ["0.0.0.0:5061"]
tls:
certificate: "/etc/siphon/tls/cert.pem"
private_key: "/etc/siphon/tls/key.pem"
method: "TLSv1_3"
# mTLS — require and verify client certs (SIP trunks with mutual auth):
verify_client: true
client_ca: "/etc/siphon/tls/client-ca.pem"
method is the minimum TLS version. TLSv1_3 here is a real 1.3-only floor —
it refuses TLS 1.2 peers on the listeners and on outbound connections siphon
dials, so check both sides can do 1.3 before hardening. TLSv1_2 (the default)
negotiates 1.2 or 1.3.
verify_client: true requires a client cert chaining to client_ca (fails closed at
startup if client_ca is missing). It applies to listen.tls and listen.wss.
4. Authenticate subscribers (script + config)¶
if not await auth.require_digest(request, realm="example.com"):
return # 401/407 challenge already sent
user = request.auth_user # the authenticated username afterwards
The auth.backend can be static, http (REST credential lookup), database, or
diameter_cx (IMS HSS). For REGISTER-time account-takeover protection, set
registrar.enforce_auth_aor_match: true so a subscriber can't bind a Contact under
someone else's AoR.
5. Verify caller ID — STIR/SHAKEN (script)¶
Sign on egress, verify on ingress at a trunk edge:
from siphon import proxy, stir, log
@proxy.on_request("INVITE")
def on_invite(request):
if request.source_ip_in(["203.0.113.0/24"]): # inbound from a peer
result = stir.verify(request)
if result.verstat == "TN-Validation-Failed":
request.reply(438, "Invalid Identity Header") # RFC 8224 §6.2.2
return
stir.apply_verstat(request, result) # convey downstream
else: # outbound
origid = stir.sign(request, attestation="A")
request.record_route()
request.relay()
Needs a stir: block with signing + verification configured.
The source_ip_in([...]) above hardcodes the peer's CIDR. If that peer is already
a gateway group (a trunk you health-probe), test membership by group name instead
so you never maintain two copies of the address list — see the next section.
5.5. Direction & trust — from_gateway¶
request.from_gateway("group") (and call.from_gateway("group") in a B2BUA) returns
True when the message's source IP is one of the resolved addresses of the named
gateway group. It's SIPhon's equivalent of Kamailio ds_is_from_list() /
OpenSIPS ds_is_in_list() — a routing-direction predicate that replaces hardcoded
source CIDRs with the trunk list you already maintain under gateway.groups.
from siphon import proxy, gateway
@proxy.on_request("INVITE")
def route(request):
if request.from_gateway("teams"):
# Inbound leg from Microsoft Teams — trust it, forward to the PBX.
request.relay("sip:pbx.internal:5060")
else:
# Outbound leg from the PBX — send to Teams.
request.relay(gateway.select("teams").uri)
It matches on IP only (source port ignored) against every resolved address in
the group, so a hostname that round-robins across many IPs — Teams'
sip/sip2/sip3.pstnhub.microsoft.com, a carrier's rotating trunk — matches on any
of them. The member set is cached and refreshed on the health-probe cycle, so the
predicate never resolves DNS on the request path.
Trustworthy on TCP/TLS/WS/WSS, a hint on UDP
On connection-oriented transports the source IP is verified by the handshake, so
from_gateway is a sound authorization signal. On UDP the source IP is
spoofable — treat from_gateway there as a best-effort direction hint, and
gate real trust decisions on TLS/mTLS or digest/AKA auth.
6. IMS access security — IPsec (Gm)¶
For a P-CSCF, SIPhon does full 3GPP TS 33.203 sec-agree: parse Security-Client,
run AKA, install kernel IPsec SAs, and route MT requests back over the flow. It's a
substantial flow — see examples/ims_pcscf.py
and the ipsec: config block. The SA lifetime tracks the registration lifetime
automatically.
A re-REGISTER or de-REGISTER over the SA is not challenged again, and that
rests on a chain of trust worth knowing. AKA nonces are single-use, so the UE
cannot reuse its old Authorization. Instead the P-CSCF calls
auth.stamp_integrity_protected(request) on every REGISTER it relays, which
writes integrity-protected="yes" into the Authorization header only when
the REGISTER came over an SA negotiated for that header's IMPI, and "no"
otherwise, whatever the UE sent. The S-CSCF calls
auth.verify_integrity_protected(request) before its AKA challenge and skips
the challenge only when the header says protected and that IMPI is the one
that registered the To IMPU (contact.auth_user). The first check stops a UE
with a valid SA of its own from claiming another subscriber's IMPI; the second
stops it de-registering someone else's IMPU. The S-CSCF takes the header on the
P-CSCF's word, so only use verify_integrity_protected behind a P-CSCF that
stamps every REGISTER.
When SIPhon terminates the Gm hop as a B2BUA, an INVITE that requires sec-agree
(in Require or Proxy-Require) is checked before your script runs, per RFC 3329
§2.3.1. It is answered 494 Security Agreement Required unless it arrived on a
protected port over an active SA and its Security-Verify mirrors the
Security-Server your script put on the 401 for the REGISTER that set that SA up.
SIPhon records that value on the SA when it relays the 401, and compares the two
parameter for parameter, q, prot and mod included: names and values ignore
case (a quoted string must match exactly), whitespace is ignored, and parameters
and entries can come in any order. PendingSA.refresh() carries the value over to
the re-keyed SA, and the 401 for the re-key records it again. An SA with nothing
recorded (installed before an upgrade, or set up from a 401 SIPhon generated itself
rather than relayed) is checked on what the SA holds instead: the ipsec-3gpp
algorithms, SPIs and protected ports. A 494 over an SA carries its recorded
Security-Server, or one built from the SA.
A 494 to an unprotected request has no SA to name. Its Security-Server lists what
SIPhon supports, one ipsec-3gpp line per transform with alg and ealg only:
Security-Server: ipsec-3gpp; alg=hmac-sha-1-96; ealg=null
Security-Server: ipsec-3gpp; alg=hmac-md5-96; ealg=null
Security-Server: ipsec-3gpp; alg=hmac-sha-256-128; ealg=null
Security-Server: ipsec-3gpp; alg=hmac-sha-1-96; ealg=aes-cbc
Security-Server: ipsec-3gpp; alg=hmac-md5-96; ealg=aes-cbc
Security-Server: ipsec-3gpp; alg=hmac-sha-256-128; ealg=aes-cbc
hmac-sha-256-128 in that list is a SIPhon extension, not a 3GPP transform. The
transform is HMAC-SHA-256-128 per RFC 4868, but its 256-bit key comes from
SIPhon's own expansion, HMAC-SHA-256(IK, "ipsec-int-sha256-128"): no 3GPP
release we have checked lists hmac-sha-256-128 in Annex H, whose key expansion
lives in Annex I. The label is arbitrary, so two implementations each inventing
one would not interoperate. Use it only where both ends are SIPhon. Note also
that Rel-13 Annex H drops hmac-md5-96 and adds aes-gmac and aes-gcm, which
SIPhon does not implement.
TS 33.203 Annex H makes spi-c, spi-s, port-c and port-s mandatory for
ipsec-3gpp. Leaving them out is deliberate: RFC 3329 §2.3.1 asks the 494 for the
server's list of supported mechanisms, and before an SA exists there are no SPIs or
ports to put in it. A UE gets the full syntax on the 401 to its REGISTER. The same
list goes on the 494 to a call your script made require sec-agree without a
verified agreement. The agreement ends at SIPhon: the
B-leg INVITE carries no sec-agree in Require or Proxy-Require and no
Security-Verify or Security-Client, unless your script sets them for an agreement
of the B-leg's own.
Checklist¶
- [ ]
security.failed_auth_ban+scanner_blockon, infra intrusted_cidrs - [ ] Behind a terminating front:
proxy_protocol.fromnames that front and only that front (it does not inherittrusted_cidrs) - [ ]
proxy.sanity_check()on out-of-dialog requests, silent-drop failures - [ ] TLS (and mTLS for trunks); subscriber-facing access over TLS/WSS
- [ ] Digest auth on REGISTER (+
enforce_auth_aor_match) - [ ] STIR/SHAKEN at PSTN edges; IPsec at IMS Gm
- [ ] Alert on the security metrics (see Monitoring)
See also¶
- Real example:
examples/stir_shaken.py,examples/ims_pcscf.py. - Reference config:
siphon.yaml.