Blog · · 5 min read

What your users notice when the SIP edge is done right

Nobody rings the helpdesk to say the phones worked today. A well-built SIP edge shows up mostly as complaints you don't get. Here is where that quiet comes from, and which parts of Kamailio produce it.

By the Kamailio Support engineers.

Phones that stay registered

Here is what users see when this goes wrong. A desk phone shows "No service" after lunch, or a softphone looks online but misses inbound calls. Usually the cause is NAT. The office router has quietly closed the pinhole, so the contact Kamailio stored points at a port that is no longer open.

The nathelper module deals with this. nat_uac_test() spots clients behind NAT, fix_nated_register() records the address the request really came from, and keepalives (nat_keepalive_interval, or SIP OPTIONS pings) hold the pinhole open. usrloc can also write registrations to the database (db_mode), so after a restart the proxy still knows where every phone is and doesn't have to wait for each one to register again.

Registration is also the first thing scanners attack. Every REGISTER should pass digest authentication with auth_check(), and repeated failures should be blocked using pike or htable counters. Treat those limits as something you keep looking after: the scanners keep changing, so read the logs and adjust the rules. Our guide on friendly-scanner and SIP brute-force attacks covers the details.

Calls that connect fast

The complaint here is post-dial delay: someone dials, hears several seconds of silence, and often hangs up before it rings. A common cause is the proxy sending the call to a gateway or PBX that has died, then waiting for a timeout before it tries the next one.

The dispatcher module avoids this by checking destinations ahead of time. When ds_ping_interval is set, it sends SIP OPTIONS to every destination and marks any that stop answering as inactive, so ds_select_dst() never picks them. If a call still fails, a failure_route calls ds_next_dst() and tries the next destination straight away. The tm timers (fr_timer, fr_inv_timer) set how long Kamailio waits before giving up. If they are tuned too loosely, people sit in silence. If they are tuned too tightly, calls get abandoned while they are still ringing.

If failover doesn't behave as you expect, see dispatcher not failing over. The dispatcher documentation lists every flag and parameter.

Both people can hear each other

One-way audio is the complaint that does most damage to a platform's reputation. The call connects and one person talks to silence. The cause is nearly always the SDP: it advertises a private address that the far end cannot reach.

Kamailio carries no audio itself, so the fix goes alongside it. rtpengine, the open-source media relay from the Sipwise team, relays the media, and rtpengine_manage() in the Kamailio config rewrites the SDP so each side sends audio to an address it can actually reach. rtpengine can also translate between media formats: it can terminate DTLS-SRTP from a WebRTC browser and hand plain RTP to an internal PBX, or encrypt media for calls leaving the building. The guide on one-way audio through Kamailio and rtpengine goes through the usual causes.

Keep rtpengine's control socket reachable from Kamailio only, never from the internet. Open its media port range in the firewall, and nothing else.

Calls that survive a server failure

Done well, a server failure is something your users never hear about. What is involved depends on which server fails.

  • A FreeSWITCH or PBX node dies: the dispatcher health checks mark it inactive and new calls go to the nodes that are still up. Calls already in progress on the dead node are harder to save. FreeSWITCH has call recovery (sofia recover), but it needs planning and testing before you rely on it.
  • A Kamailio node dies: run two of them with a floating IP (keepalived is the usual choice). Replicate registrations with dmq and dmq_usrloc, and the standby already knows every phone when it takes over.
  • Calls already in progress: Kamailio doesn't carry the audio, so on a well-built platform the media keeps flowing through rtpengine while the signalling address moves. The dialog module can share call state over DMQ, so call counts and limits stay accurate.

Test failover on a schedule, not only when you first build it. A standby that has never taken over is only an assumption.

Caller ID that people trust

This matters most to the people you call. If spoofed numbers and nuisance calls have taught them to ignore unknown callers, your outbound calls go unanswered. In the UK, Ofcom requires calls to carry a valid, dialable CLI, and networks are expected to stop calls that don't. The SIP edge is where that rule is enforced.

In practice: authenticate the caller, discard any identity headers they sent about themselves, check the number against the ones that account is allowed to present, then assert it yourself. Where carriers use STIR/SHAKEN, Kamailio's stirshaken and secsipid modules can sign and verify the Identity header.

kamailio.cfg (trimmed)
modparam("htable", "db_url", DBURL)
modparam("htable", "htable", "cli=>size=10;dbtable=allowed_cli;")

route[CALLER_ID] {
    # Callers must authenticate before anything else
    if (!auth_check("$fd", "subscriber", "1")) {
        auth_challenge("$fd", "0");
        exit;
    }
    consume_credentials();

    # Never pass on identity a caller asserts about itself
    remove_hf("P-Asserted-Identity");

    # Only present numbers this account may use
    if ($sht(cli=>$au::$fU) == $null) {
        sl_send_reply("403", "Caller ID not allowed");
        exit;
    }
    append_hf("P-Asserted-Identity: <sip:$fU@$fd>\r\n");
}

Limit fraud at the same point. Track calls in a dialog profile with set_dlg_profile() and check get_profile_size() before a new call goes out. Then a stolen password can't run a hundred international calls overnight. As fraud patterns change, review the limits and your list of blocked destinations.

What the ops team notices

The on-call phone is quieter. When something does break, there is evidence to work from: SIP traces sent over HEP with siptrace, metrics exported by xhttp_prom and graphed in Grafana, and logs that show which rule rejected a call. The platform runs a supported Kamailio release and gets patched, and the security rules are reviewed regularly, not left alone once they seem to work.

Most of our Kamailio consultancy and support work is building and maintaining exactly this. If you run Kamailio yourself, the project's documentation and the sr-users mailing list are the best places to start.

Comments

No comments yet.

Add a comment

One of our engineers reads every comment before it appears. Please don't post passwords, IP addresses or other private details from your platform.

We publish your name and comment with this post. To have a comment removed, email [email protected].

Work with us

Want this on your platform?

Three senior engineers who've run Kamailio and FreeSWITCH in production since 2005. You talk to the person doing the work, and we'll sign your NDA before you share anything sensitive. See pricing, or read about our Kamailio support.

Call us Email