ST 2110-30 & AES67 – Professional Audio over IP


BL Study Plan2110 Topo

What you will learn on this page


This lesson explains how the 2110-30 & the AES67 Standards fits into the SMPTE ST 2110 stack.
ST 2110-30 & AES67 – Professional Audio over IP
ST 2110-30 vs AES67: Compatibility vs Full Compliance
AES67 Operational Levels (Profiles)
What a 2110 Engineer Must Know About ST 2110-30 Audio
Common Troubleshooting Scenarios for ST 2110-30
Historical note


ST 2110-30 defines uncompressed PCM audio transport within the ST 2110 suite. It is built directly on top of AES67, the Audio Engineering Society’s interoperability standard.

ST 2110-30 Audio Flow
ST 2110-30 Audio Flow – Uncompressed PCM over RTP Multicast

ST 2110-30 & AES67 – Professional Audio over IP

ST 2110-30 is built directly on AES67. While AES67 is a general interoperability standard, ST 2110-30 adds broadcast-specific rules and recommendations.

Feature AES67 ST 2110-30
Base Standard AES67 – General professional audio-over-IP interoperability standard. AES67 + SMPTE-specific constraints for live broadcast production.
Sampling Rate Supports 44.1, 48, 88.2, and 96 kHz. Primarily 48 kHz (96 kHz allowed). 44.1 kHz is discouraged in broadcast plants.
Packet Time (ptime) Flexible: 0.125 ms, 0.25 ms, 1 ms, 4 ms. Strongly recommends 1 ms. This is the sweet spot for most 2110 facilities.
Channels per Stream Up to 8 channels (very flexible). Typically 1 to 8 channels per stream. Often organized as mono or stereo pairs.
Timing & Synchronization Bottom Line   Requires IEEE 1588 PTP. Must use the ST 2059 PTP profile (same as video). Same PTP domain as video is mandatory for lip-sync.
Interoperability Broad compatibility across many manufacturers and systems. AES67 compatible by design, plus additional broadcast operational rules for reliability.

Detailed Explanation of Each Row

Base Standard
AES67 is the foundational interoperability spec. ST 2110-30 takes AES67 and adds stricter rules required for live television and broadcast production environments.
Sampling Rate
While AES67 allows multiple rates, ST 2110-30 facilities almost always standardize on 48 kHz to match video frame rates and simplify synchronization.
Packet Time (ptime)
Bottom Line   Shorter packet times reduce latency but increase network overhead. 1 ms is the recommended default in 2110-30 because it balances latency, network load, and jitter tolerance.
Channels per Stream
Both standards support up to 8 channels per flow. In practice, 2110-30 engineers usually send audio as mono or stereo pairs for easier routing and monitoring.
Timing & Synchronization
This is the most important difference. ST 2110-30 requires all audio to use the same PTP Grandmaster and domain as video. This guarantees automatic lip-sync.
Interoperability
ST 2110-30 devices are always AES67 compatible, but not all AES67 devices fully comply with 2110-30’s stricter timing and operational recommendations.
Bottom Line for 2110 Engineers:
Treat ST 2110-30 as “AES67 with broadcast discipline.” Master 48 kHz, 1 ms packet time, and tight PTP synchronization with video. Once these are correct, audio integration becomes reliable and lip-sync is automatic.

ST 2110-30 vs AES67: Compatibility vs Full Compliance

ST 2110-30 devices are always AES67 compatible, but the reverse is not always true. This is one of the most important nuances a 2110 engineer must understand.

The Relationship

ST 2110-30 is essentially AES67 with additional broadcast-specific rules. Think of AES67 as the general foundation, and ST 2110-30 as a stricter “broadcast profile” built on top of it.

Key Differences – What Makes ST 2110-30 More Demanding

Aspect AES67 (General) ST 2110-30 (Broadcast)
Sampling Rate 44.1 kHz, 48 kHz, 88.2 kHz, 96 kHz Strongly recommends 48 kHz only
Packet Time (ptime) 0.125, 0.25, 1, or 4 ms (very flexible) Recommends 1 ms as the primary operating mode
PTP Timing Profile Any IEEE 1588 PTP implementation Must use ST 2059-2 PTP profile
Synchronization with Video Not required Mandatory – must share same PTP domain and clock as video for lip-sync
Operational Discipline General purpose audio networking Designed for live, on-air broadcast reliability and redundancy

Practical Implications for 2110 Engineers

Rule of Thumb:
All ST 2110-30 devices are AES67 compatible.
Not all AES67 devices are suitable for serious ST 2110-30 broadcast deployments.

AES67 Operational Levels (Profiles)

Bottom Line   To ensure reliable cross-vendor interoperability, AES67 defines three Operational Levels (also called profiles). These levels specify exactly which features a device must support.

AES67 Operational Levels Comparison

Feature Level A Level B Level C
Sampling Rate 48 kHz only 48 kHz 48 kHz
Packet Time (ptime) 1 ms only 0.125, 0.25, 1 ms 0.125, 0.25, 1, 4 ms
Max Channels per Stream 8 8 8
Recommended For Most Broadcast Facilities
(Default for ST 2110-30)
Low-latency applications Advanced / specialized use

Why These Levels Matter in ST 2110-30

Bottom Line:
For most ST 2110 deployments, standardize on AES67 Level A (48 kHz, 1 ms packet time). This ensures maximum compatibility and reliable performance across vendors.

What a 2110 Engineer Must Know About ST 2110-30 Audio

Successfully implementing and troubleshooting ST 2110-30 audio requires deep understanding of several critical areas.

ST 2110-30 Audio Flow
ST 2110-30 Audio Flow – Uncompressed PCM over RTP Multicast

ST 2110-30 Audio Flow Explained

Detailed Breakdown of Key Elements in the Diagram

SDP (Session Description Protocol) + "1ms Packet Time"
The SDP is the "contract" or configuration file that tells receivers how to decode the audio stream. The label "1ms Packet Time" refers to the ptime parameter in the SDP. This tells the receiver that audio packets are sent every 1 millisecond.

Why it matters: 1 ms is the recommended default in ST 2110-30. It provides a good balance between low latency and manageable network load. Shorter packet times increase overhead; longer ones increase latency.
RTP Timestamps + PTP Timestamp
Bottom Line   Each RTP packet contains a timestamp that tells the receiver exactly when that audio sample should be played. In ST 2110-30, this timestamp is derived from the facility’s shared PTP clock (not the local computer clock).

Why it matters: Because both video (2110-20) and audio (2110-30) use the same PTP time reference, the receiver can perfectly align audio with video — resulting in automatic, stable lip-sync across the entire system.
Summary:
SDP tells receivers how the stream is formatted (including packet timing), while RTP Timestamps (derived from PTP) tell receivers when to play each sample — enabling precise synchronization with video.

Key Knowledge Areas

1. PTP Synchronization
Audio must use the exact same PTP Grandmaster and domain as video. This is non-negotiable in ST 2110. If audio and video are in different PTP domains, lip-sync will fail. Best practice: Use Domain 127 for the entire facility and verify all devices report the same Grandmaster with sub-microsecond offset.
2. Packet Timing (ptime)
The recommended default in ST 2110-30 is 1 ms packet time. This provides the best balance between latency and network efficiency. Shorter packet times increase network load significantly, while longer times increase latency.
3. SDP Configuration
Bottom Line   Accurate SDP files are critical. Key parameters you must get right:
  • ptime: Usually 1 (1 ms)
  • channels: Number of audio channels in the stream
  • clock-rate: Almost always 48000
Incorrect SDP is one of the most common causes of audio not appearing at the receiver.
4. QoS Marking
Audio traffic is typically marked as CS5 or AF41 — high priority, but usually one step below video. PTP timing packets remain at CS6 (highest priority).
5. Channel Grouping
Understand how to map audio efficiently: mono streams, stereo pairs, or multi-channel groups (up to 8 channels per flow is common).
6. NMOS Control (IS-04 & IS-05)
Use IS-04 for device discovery and IS-05 for dynamic connection management of audio flows.
ST 2110-30 PTP Synchronization
How ST 2110-30 Audio Stays Synchronized with Video via Shared PTP

Common Troubleshooting Scenarios for ST 2110-30

Pro Tip:
Always verify PTP lock and offset first (before checking audio). Most audio problems in ST 2110 systems are actually timing or network-related, not audio-specific.
ST 2110-30 PTP Synchronization
How ST 2110-30 Audio Stays Synchronized with Video via PTP
Bottom Line:
ST 2110-30 is AES67 with broadcast-specific rules. Master PTP timing, 1 ms packet time, and clean SDP files. Once audio is locked to the same PTP clock as video, lip-sync becomes automatic and rock-solid.

An historical note for those of you old enough to remember the original AES standart

AES3 vs AES67 – What’s the Difference?
AES3 and AES67 are both standards for digital audio, but they represent two completely different technologies and eras.

AES3 (also known as AES/EBU) is the traditional, long-established standard (introduced in 1985). It carries uncompressed stereo digital audio over physical cables — typically XLR or BNC. It embeds the clock within the audio signal itself and is limited to point-to-point connections with a maximum practical distance of about 100–300 meters.

AES67 is a modern Audio-over-IP standard published in 2013. It transports uncompressed PCM audio over Ethernet/IP networks using RTP multicast. It relies on external PTP (Precision Time Protocol) for synchronization instead of embedding the clock in the signal. This allows flexible any-to-any routing across a facility or even across campuses.

Key Differences:

Bottom Line:
AES3 is the old "copper" digital audio standard. AES67 is the modern network-based standard. In ST 2110 systems, ST 2110-30 is essentially AES67 with additional broadcast-specific rules for timing and reliability.
UPDATED
5/24/26
V260525-1.0