ST 2110-30 defines uncompressed PCM audio transport within the ST 2110 suite. It is built directly on top of AES67, the Audio Engineering Society’s interoperability standard.
ST 2110-30 Audio Flow – Uncompressed PCM over RTP Multicast
ST 2110-30 & AES67 – Professional Audio over IP
ST 2110-30 is built directly on AES67. While AES67 is a general interoperability standard, ST 2110-30 adds broadcast-specific rules and recommendations.
Feature
AES67
ST 2110-30
Base Standard
AES67 – General professional audio-over-IP interoperability standard.
AES67 + SMPTE-specific constraints for live broadcast production.
Sampling Rate
Supports 44.1, 48, 88.2, and 96 kHz.
Primarily 48 kHz (96 kHz allowed). 44.1 kHz is discouraged in broadcast plants.
Packet Time (ptime)
Flexible: 0.125 ms, 0.25 ms, 1 ms, 4 ms.
Strongly recommends 1 ms. This is the sweet spot for most 2110 facilities.
Channels per Stream
Up to 8 channels (very flexible).
Typically 1 to 8 channels per stream. Often organized as mono or stereo pairs.
Timing & Synchronization
Requires IEEE 1588 PTP.
Must use the ST 2059 PTP profile (same as video). Same PTP domain as video is mandatory for lip-sync.
Interoperability
Broad compatibility across many manufacturers and systems.
AES67 compatible by design, plus additional broadcast operational rules for reliability.
Detailed Explanation of Each Row
Base Standard
AES67 is the foundational interoperability spec. ST 2110-30 takes AES67 and adds stricter rules required for live television and broadcast production environments.
Sampling Rate
While AES67 allows multiple rates, ST 2110-30 facilities almost always standardize on 48 kHz to match video frame rates and simplify synchronization.
Packet Time (ptime)
Shorter packet times reduce latency but increase network overhead. 1 ms is the recommended default in 2110-30 because it balances latency, network load, and jitter tolerance.
Channels per Stream
Both standards support up to 8 channels per flow. In practice, 2110-30 engineers usually send audio as mono or stereo pairs for easier routing and monitoring.
Timing & Synchronization
This is the most important difference. ST 2110-30 requires all audio to use the same PTP Grandmaster and domain as video. This guarantees automatic lip-sync.
Interoperability
ST 2110-30 devices are always AES67 compatible, but not all AES67 devices fully comply with 2110-30’s stricter timing and operational recommendations.
Bottom Line for 2110 Engineers:
Treat ST 2110-30 as “AES67 with broadcast discipline.” Master 48 kHz, 1 ms packet time, and tight PTP synchronization with video. Once these are correct, audio integration becomes reliable and lip-sync is automatic.
ST 2110-30 vs AES67: Compatibility vs Full Compliance
ST 2110-30 devices are always AES67 compatible, but the reverse is not always true. This is one of the most important nuances a 2110 engineer must understand.
The Relationship
ST 2110-30 is essentially AES67 with additional broadcast-specific rules. Think of AES67 as the general foundation, and ST 2110-30 as a stricter “broadcast profile” built on top of it.
Key Differences – What Makes ST 2110-30 More Demanding
Aspect
AES67 (General)
ST 2110-30 (Broadcast)
Sampling Rate
44.1 kHz, 48 kHz, 88.2 kHz, 96 kHz
Strongly recommends 48 kHz only
Packet Time (ptime)
0.125, 0.25, 1, or 4 ms (very flexible)
Recommends 1 ms as the primary operating mode
PTP Timing Profile
Any IEEE 1588 PTP implementation
Must use ST 2059-2 PTP profile
Synchronization with Video
Not required
Mandatory – must share same PTP domain and clock as video for lip-sync
Operational Discipline
General purpose audio networking
Designed for live, on-air broadcast reliability and redundancy
Practical Implications for 2110 Engineers
You can connect an AES67 device to a 2110-30 system and it will usually work at a basic level.
However, the AES67 device may not follow ST 2110-30’s stricter timing rules, leading to:
Lip-sync problems with video
Higher susceptibility to jitter
Inconsistent behavior during redundancy (2022-7) events
Full ST 2110-30 compliance ensures seamless integration with video, reliable PTP lock, and predictable performance in live broadcast environments.
Rule of Thumb:
All ST 2110-30 devices are AES67 compatible.
Not all AES67 devices are suitable for serious ST 2110-30 broadcast deployments.
AES67 Operational Levels (Profiles)
To ensure reliable cross-vendor interoperability, AES67 defines three Operational Levels (also called profiles). These levels specify exactly which features a device must support.
AES67 Operational Levels Comparison
Feature
Level A
Level B
Level C
Sampling Rate
48 kHz only
48 kHz
48 kHz
Packet Time (ptime)
1 ms only
0.125, 0.25, 1 ms
0.125, 0.25, 1, 4 ms
Max Channels per Stream
8
8
8
Recommended For
Most Broadcast Facilities
(Default for ST 2110-30)
Low-latency applications
Advanced / specialized use
Why These Levels Matter in ST 2110-30
Level A is the recommended and most commonly used level in broadcast facilities. It offers the best balance of performance, network efficiency, and reliability.
ST 2110-30 strongly encourages devices to support at least Level A.
Higher levels (B and C) add support for lower latency packet times, but they increase network load and complexity.
When purchasing equipment, always verify which AES67 levels are supported — especially if you plan to use it in a 2110-30 environment.
Bottom Line:
For most ST 2110 deployments, standardize on AES67 Level A (48 kHz, 1 ms packet time).
This ensures maximum compatibility and reliable performance across vendors.
What a 2110 Engineer Must Know About ST 2110-30 Audio
Successfully implementing and troubleshooting ST 2110-30 audio requires deep understanding of several critical areas.
ST 2110-30 Audio Flow – Uncompressed PCM over RTP Multicast
The SDP is the "contract" or configuration file that tells receivers how to decode the audio stream.
The label "1ms Packet Time" refers to the ptime parameter in the SDP.
This tells the receiver that audio packets are sent every 1 millisecond.
Why it matters: 1 ms is the recommended default in ST 2110-30. It provides a good balance between low latency and manageable network load.
Shorter packet times increase overhead; longer ones increase latency.
RTP Timestamps + PTP Timestamp
Each RTP packet contains a timestamp that tells the receiver exactly when that audio sample should be played.
In ST 2110-30, this timestamp is derived from the facility’s shared PTP clock (not the local computer clock).
Why it matters: Because both video (2110-20) and audio (2110-30) use the same PTP time reference,
the receiver can perfectly align audio with video — resulting in automatic, stable lip-sync across the entire system.
Summary:
SDP tells receivers how the stream is formatted (including packet timing), while RTP Timestamps (derived from PTP) tell receivers when to play each sample — enabling precise synchronization with video.
Key Knowledge Areas
1. PTP Synchronization
Audio must use the exact same PTP Grandmaster and domain as video.
This is non-negotiable in ST 2110. If audio and video are in different PTP domains, lip-sync will fail.
Best practice: Use Domain 127 for the entire facility and verify all devices report the same Grandmaster with sub-microsecond offset.
2. Packet Timing (ptime)
The recommended default in ST 2110-30 is 1 ms packet time.
This provides the best balance between latency and network efficiency.
Shorter packet times increase network load significantly, while longer times increase latency.
3. SDP Configuration
Accurate SDP files are critical.
Key parameters you must get right:
ptime: Usually 1 (1 ms)
channels: Number of audio channels in the stream
clock-rate: Almost always 48000
Incorrect SDP is one of the most common causes of audio not appearing at the receiver.
4. QoS Marking
Audio traffic is typically marked as CS5 or AF41 — high priority, but usually one step below video.
PTP timing packets remain at CS6 (highest priority).
5. Channel Grouping
Understand how to map audio efficiently: mono streams, stereo pairs, or multi-channel groups (up to 8 channels per flow is common).
6. NMOS Control (IS-04 & IS-05)
Use IS-04 for device discovery and IS-05 for dynamic connection management of audio flows.
How ST 2110-30 Audio Stays Synchronized with Video via Shared PTP
Common Troubleshooting Scenarios for ST 2110-30
No audio at receiver → Wrong SDP (especially ptime, channels, or multicast address), IGMP subscription issue, or firewall blocking.
Audio clicks, pops, or dropouts → Packet loss, jitter, or QoS misconfiguration.
Switch queue starvation → Occurs when a high-priority queue (such as one carrying PTP timing packets marked CS6) consumes so much of the switch’s resources that lower-priority queues (like those carrying video or audio media) are starved of bandwidth and get delayed or dropped.
Lip-sync problems with video → Different PTP domains, one device not locked to Grandmaster, or mismatched clock rates.
High latency → Incorrect packet time (ptime) or excessive buffering on the receiver.
Intermittent audio → IGMP Querier not working, PIM-SSM misconfiguration (if routed), or multicast group limits exceeded.
2022-7 not hitless on audio → One Red/Blue path has different PTP behavior or higher jitter.
Pro Tip:
Always verify PTP lock and offset first (before checking audio).
Most audio problems in ST 2110 systems are actually timing or network-related, not audio-specific.
How ST 2110-30 Audio Stays Synchronized with Video via PTP
Bottom Line:
ST 2110-30 is AES67 with broadcast-specific rules. Master PTP timing, 1 ms packet time, and clean SDP files.
Once audio is locked to the same PTP clock as video, lip-sync becomes automatic and rock-solid.
An historical note for those of you old enough to remember the original AES standart
AES3 vs AES67 – What’s the Difference?
AES3 and AES67 are both standards for digital audio, but they represent two completely different technologies and eras.
AES3 (also known as AES/EBU) is the traditional, long-established standard (introduced in 1985). It carries uncompressed stereo digital audio over physical cables — typically XLR or BNC. It embeds the clock within the audio signal itself and is limited to point-to-point connections with a maximum practical distance of about 100–300 meters.
AES67 is a modern Audio-over-IP standard published in 2013. It transports uncompressed PCM audio over Ethernet/IP networks using RTP multicast. It relies on external PTP (Precision Time Protocol) for synchronization instead of embedding the clock in the signal. This allows flexible any-to-any routing across a facility or even across campuses.
Key Differences:
Transport: AES3 uses physical cables; AES67 uses IP networks.
Routing: AES3 is point-to-point (requires patch panels or routers); AES67 allows dynamic network routing.
Timing: AES3 embeds clock in the signal; AES67 uses PTP.
Distance & Scalability: AES3 is limited by cable length; AES67 can cover entire facilities and beyond.
Channels: AES3 carries 2 channels per link; AES67 can carry up to 8 channels per stream.
Role in ST 2110: AES3 is legacy (often converted via gateways); AES67 is the foundation of ST 2110-30.
Bottom Line:
AES3 is the old "copper" digital audio standard. AES67 is the modern network-based standard. In ST 2110 systems, ST 2110-30 is essentially AES67 with additional broadcast-specific rules for timing and reliability.