Request for Comments: 4348 January 2006
Category: Standards Track
Real-Time Transport Protocol (RTP) Payload Format for the
Variable-Rate Multimode Wideband (VMR-WB) Audio Codec
Status of This Memo
This document specifies an Internet standards track protocol for the
Internet community, and requests discussion and suggestions for
improvements. Please refer to the current edition of the "Internet
Official Protocol Standards" (STD 1) for the standardization state
and status of this protocol. Distribution of this memo is unlimited.
Copyright Notice
Copyright (C) The Internet Society (2006).
Abstract
This document specifies a real-time transport protocol (RTP) payload
format to be used for the Variable-Rate Multimode Wideband (VMR-WB)
speech codec. The payload format is designed to be able to
interoperate with existing VMR-WB transport formats on non-IP
networks. A media type registration is included for VMR-WB RTP
payload format.
VMR-WB is a variable-rate multimode wideband speech codec that has a
number of operating modes, one of which is interoperable with AMR-WB
(i.e., RFC 3267) audio codec at certain rates. Therefore, provisions
have been made in this document to facilitate and simplify data
packet exchange between VMR-WB and AMR-WB in the interoperable mode
with no transcoding function involved.
Table of Contents
1. Introduction ....................................................3
2. Conventions and Acronyms ........................................3
3. The Variable-Rate Multimode Wideband (VMR-WB) Speech Codec ......4
3.1. Narrowband Speech Processing ...............................5
3.2. Continuous vs. Discontinuous Transmission ..................6
3.3. Support for Multi-Channel Session ..........................6
4. Robustness against Packet Loss ..................................7
4.1. Forward Error Correction (FEC) .............................7
4.2. Frame Interleaving and Multi-Frame Encapsulation ...........8
5. VMR-WB Voice over IP Scenarios ..................................9
5.1. IP Terminal to IP Terminal .................................9
5.2. GW to IP Terminal .........................................10
5.3. GW to GW (between VMR-WB- and AMR-WB-Enabled Terminals) ...10
5.4. GW to GW (between Two VMR-WB-Enabled Terminals) ...........11
6. VMR-WB RTP Payload Formats .....................................12
6.1. RTP Header Usage ..........................................13
6.2. Header-Free Payload Format ................................14
6.3. Octet-Aligned Payload Format ..............................15
6.3.1. Payload Structure ..................................15
6.3.2. The Payload Header .................................15
6.3.3. The Payload Table of Contents ......................18
6.3.4. Speech Data ........................................20
6.3.5. Payload Example: Basic Single Channel
Payload Carrying Multiple Frames ...................21
6.4. Implementation Considerations .............................22
6.4.1. Decoding Validation and Provision for Lost
or Late Packets ....................................22
7. Congestion Control .............................................23
8. Security Considerations ........................................23
8.1. Confidentiality ...........................................24
8.2. Authentication and Integrity ..............................24
9. Payload Format Parameters ......................................24
9.1. VMR-WB RTP Payload MIME Registration ......................25
9.2. Mapping MIME Parameters into SDP ..........................27
9.3. Offer-Answer Model Considerations .........................28
10. IANA Considerations ...........................................29
11. Acknowledgements ..............................................29
12. References ....................................................30
12.1. Normative References .....................................30
12.2. Informative References ...................................30
1. Introduction
This document specifies the payload format for packetization of VMR-
WB-encoded speech signals into the Real-time Transport Protocol (RTP)
[3]. The VMR-WB payload formats support transmission of single and
multiple channels, frame interleaving, multiple frames per payload,
header-free payload, the use of mode switching, and interoperation
with existing VMR-WB transport formats on non-IP networks, as
described in Section 3.
The payload format is described in Section 6. The VMR-WB file format
(i.e., for transport of VMR-WB speech data in storage mode
applications such as email) is specified in [7]. In Section 9, a
media type registration for VMR-WB RTP payload format is provided.
Since VMR-WB is interoperable with AMR-WB at certain rates, an
attempt has been made throughout this document to maximize the
similarities with RFC 3267 while optimizing the payload format for
the non-interoperable modes of the VMR-WB codec.
2. Conventions and Acronyms
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
"SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this
document are to be interpreted as described in RFC2119 [2].
The following acronyms are used in this document:
3GPP - The Third Generation Partnership Project
3GPP2 - The Third Generation Partnership Project 2
CDMA - Code Division Multiple Access
WCDMA - Wideband Code Division Multiple Access
GSM - Global System for Mobile Communications
AMR-WB - Adaptive Multi-Rate Wideband Codec
VMR-WB - Variable-Rate Multimode Wideband Codec
CMR - Codec Mode Request
GW - Gateway
DTX - Discontinuous Transmission
FEC - Forward Error Correction
SID - Silence Descriptor
TrFO - Transcoder-Free Operation
UDP - User Datagram Protocol
RTP - Real-Time Transport Protocol
RTCP - RTP Control Protocol
MIME - Multipurpose Internet Mail Extension
SDP - Session Description Protocol
VoIP - Voice-over-IP
The term "interoperable mode" in this document refers to VMR-WB mode
3, which is interoperable with AMR-WB codec modes 0, 1, and 2.
The term "non-interoperable modes" in this document refers to VMR-WB
modes 0, 1, and 2.
The term "frame-block" is used in this document to describe the
time-synchronized set of speech frames in a multi-channel VMR-WB
session. In particular, in an N-channel session, a frame-block will
contain N speech frames, one from each of the channels, and all N
speech frames represent exactly the same time period.
3. The Variable-Rate Multimode Wideband (VMR-WB) Speech Codec
VMR-WB is the wideband speech-coding standard developed by Third
Generation Partnership Project 2 (3GPP2) for encoding/decoding
wideband/narrowband speech content in multimedia services in 3G CDMA
cellular systems [1]. VMR-WB is a source-controlled variable-rate
multimode wideband speech codec. It has a number of operating modes,
where each mode is a tradeoff between voice quality and average data
rate. The operating mode in VMR-WB (as shown in Table 2) is chosen
based on the traffic condition of the network and the desired quality
of service. The desired average data rate (ADR) in each mode is
obtained by encoding speech frames at permissible rates (as shown in
Tables 1 and 3) compliant with CDMA2000 system, depending on the
instantaneous characteristics of input speech and the maximum and
minimum rate constraints imposed by the network operator.
While VMR-WB is a native CDMA codec complying with all CDMA system
requirements, it is further interoperable with AMR-WB [4,12] at
12.65, 8.85, and 6.60 kbps. This is due to the fact that VMR-WB and
AMR-WB share the same core technology. This feature enables
Transcoder-Free (TrFO) interconnections between VMR-WB and AMR-WB
across different wireless/wireline systems (e.g., GSM/WCDMA and
CDMA2000) without use of unnecessary complex media format conversion.
Note that the concept of mode in VMR-WB is different from that of
AMR-WB where each fixed-rate AMR-WB codec mode is adapted to
prevailing channel conditions by a tradeoff between the total number
of source-coding and channel-coding bits.
VMR-WB is able to transition between various modes with no
degradation in voice quality that is attributable to the mode
switching itself. The operating mode of the VMR-WB encoder may be
switched seamlessly without prior knowledge of the decoder. Any
non-interoperable mode (i.e., VMR-WB modes 0, 1, or 2) can be chosen
depending on the traffic conditions (e.g., network congestion) and
the desired quality of service.
While in the interoperable mode (i.e., VMR-WB mode 3), mode switching
between VMR-WB modes is not allowed because there is only one AMR-WB
interoperable mode in VMR-WB. Since the AMR-WB codec may request a
mode change, depending on channel conditions, in-band data included
in VMR-WB frame structure (see Section 8 of [1] for more details) is
used during an interoperable interconnection to switch between VMR-WB
frame types 0, 1, and 2 in VMR-WB mode 3 (corresponding to AMR-WB
codec modes 0, 1, or 2).
As mentioned earlier, VMR-WB is compliant with CDMA2000 system with
the permissible encoding rates shown in Table 1.
+---------------------------+-----------------+---------------+
| Frame Type | Bits per Packet | Encoding Rate |
| | (Frame Size) | (kbps) |
+---------------------------+-----------------+---------------+
| Full-Rate | 266 | 13.3 |
| Half-Rate | 124 | 6.2 |
| Quarter-Rate | 54 | 2.7 |
| Eighth-Rate | 20 | 1.0 |
| Blank | 0 | 0 |
| Erasure | 0 | 0 |
+---------------------------+-----------------+---------------+
Table 1: CDMA2000 system permissible frame types and their
associated encoding rates
VMR-WB is robust to high percentage of frame loss and frames with
corrupted rate information. The reception of an Erasure
(SPEECH_LOST) frame type at decoder invokes the built-in frame error
concealment mechanism. The built-in frame error concealment
mechanism in VMR-WB conceals the effect of lost frames by exploiting
in-band data and the information available in the previous frames.
3.1. Narrowband Speech Processing
VMR-WB has the capability to operate with either 16000-Hz or 8000-Hz
sampled input/output speech signals in all modes of operation [1].
The VMR-WB decoder does not require a priori knowledge about the
sampling rate of the original media (i.e., speech/audio signals
sampled at 8 or 16 kHz) at the input of the encoder. The VMR-WB
decoder, by default, generates 16000-Hz wideband output regardless of
the encoder input sampling frequency. Depending on the application,
the decoder can be configured to generate 8000-Hz output, as well.
Therefore, while this specification defines a 16000-Hz RTP clock rate
for VMR-WB codec, the injection and processing of 8000-Hz narrowband
media during a session is also allowed; however, a 16000-Hz RTP clock
rate MUST always be used.
The choice of VMR-WB output sampling frequency depends on the
implementation and the audio acoustic capabilities of the receiving
side.
3.2. Continuous vs. Discontinuous Transmission
The circuit-switched operation of VMR-WB within a CDMA network
requires continuous transmission of the speech data during a
conversation. The intrinsic source-controlled variable-rate feature
of the CDMA speech codecs is required for optimal operation of the
CDMA system and interference control. However, VMR-WB has the
capability to operate in a discontinuous transmission mode for some
packet-switched applications over IP networks (e.g., VoIP), where the
number of transmitted bits and packets during silence period are
reduced to a minimum. The VMR-WB DTX operation is similar to that of
AMR-WB [4,12].
3.3. Support for Multi-Channel Session
The octet-aligned RTP payload format defined in this document
supports multi-channel audio content (e.g., a stereophonic speech
session). Although VMR-WB codec itself does not support encoding of
multi-channel audio content into a single bit stream, it can be used
to encode and decode each of the individual channels separately.
To transport the separately encoded multi-channel content, the speech
frames for all channels that are framed and encoded for the same 20
ms periods are logically collected in a frame-block.
At the session setup, out-of-band signaling must be used to indicate
the number of channels in the session and the order of the speech
frames from different channels in each frame-block. When using SDP
for signaling (see Section 9.2 for more details), the number of
channels is specified in the rtpmap attribute, and the order of
channels carried in each frame-block is implied by the number of
channels as specified in Section 4.1 in [6].
4. Robustness against Packet Loss
The octet-aligned payload format described in this document (see
Section 6 for more details) supports several features, including
forward error correction (FEC) and frame interleaving, in order to
increase robustness against lost packets.
4.1. Forward Error Correction (FEC)
The simple scheme of repetition of previously sent data is one way of
achieving FEC. Another possible scheme, which is more bandwidth
efficient, is to use payload-external FEC; e.g., RFC2733 [8], which
generates extra packets containing repair data.
The repetition method involves the simple retransmission of
previously transmitted frame-blocks together with the current frame-
block(s). This is done by using a sliding window to group the speech
frame-blocks to send in each payload. Figure 1 illustrates an
example.
In this example, each frame-block is retransmitted one time in the
following RTP payload packet. Here, f(n-2)..f(n+4) denotes a
sequence of speech frame-blocks, and p(n-1)..p(n+4) a sequence of
payload packets.
--+--------+--------+--------+--------+--------+--------+--------+--
| f(n-2) | f(n-1) | f(n) | f(n+1) | f(n+2) | f(n+3) | f(n+4) |
--+--------+--------+--------+--------+--------+--------+--------+--
<---- p(n-1) ---->
<----- p(n) ----->
<---- p(n+1) ---->
<---- p(n+2) ---->
<---- p(n+3) ---->
<---- p(n+4) ---->
Figure 1: An example of redundant transmission
The use of this approach does not require signaling at the session
setup. In other words, the speech sender can choose to use this
scheme without consulting the receiver. This is because a packet
containing redundant frames will not look different from a packet
with only new frames. The receiver may receive multiple copies or
versions of a frame for a certain timestamp if no packet is lost. If
multiple versions of the same speech frame are received, it is
RECOMMENDED that the highest rate be used by the speech decoder.
This redundancy scheme provides the same functionality as that
described in RFC 2198, "RTP Payload for Redundant Audio Data" [10].
In most cases, the mechanism in this payload format is more efficient
and simpler than requiring both endpoints to support RFC 2198. If
the spread in time required between the primary and redundant
encodings is larger than 5 frame times, the bandwidth overhead of RFC
2198 will be lower.
The sender is responsible for selecting an appropriate amount of
redundancy based on feedback about the channel (e.g., in RTCP
receiver reports) or network traffic. A sender SHOULD NOT base
selection of FEC on the CMR, as this parameter most probably was set
based on non-IP information. The sender is also responsible for
avoiding congestion, which may be aggravated by redundant
transmission (see Section 7).
4.2. Frame Interleaving and Multi-Frame Encapsulation
To decrease protocol overhead, the octet-aligned payload format,
described in Section 6, allows several speech frame-blocks to be
encapsulated into a single RTP packet. One of the drawbacks of this
approach is that in case of packet loss several consecutive speech
frame-blocks are lost, which usually causes clearly audible
distortion in the reconstructed speech.
Interleaving of frame-blocks can improve the speech quality in such
cases by distributing the consecutive losses into a series of single
frame-block losses. However, interleaving and bundling several
frame-blocks per payload will also increase end-to-end delay and is
therefore not appropriate for all types of applications. Streaming
applications will most likely be able to exploit interleaving to
improve speech quality in lossy transmission conditions.
The octet-aligned payload format supports the use of frame
interleaving as an option. For the encoder (speech sender) to use
frame interleaving in its outbound RTP packets for a given session,
the decoder (speech receiver) needs to indicate its support via out-
of-band means (see Section 9).
5. VMR-WB Voice over IP Scenarios
5.1. IP Terminal to IP Terminal
The primary scenario for this payload format is IP end-to-end between
two terminals incorporating VMR-WB codec, as shown in Figure 2.
Nevertheless, this scenario can be generalized to an interoperable
interconnection between VMR-WB-enabled and AMR-WB-enabled IP
terminals using the offer-answer model described in Section 9.3.
This payload format is expected to be useful for both conversational
and streaming services.
+----------+ +----------+
| | | |
| TERMINAL |<----------------------->| TERMINAL |
| | VMR-WB/RTP/UDP/IP | |
+----------+ +----------+
(or AMR-WB/RTP/UDP/IP)
Figure 2: IP terminal to IP terminal
A conversational service puts requirements on the payload format.
Low delay is a very important factor, i.e., fewer speech frame-blocks
per payload packet. Low overhead is also required when the payload
format traverses across low bandwidth links, especially if the
frequency of packets will be high.
Streaming service has less strict real-time requirements and
therefore can use a larger number of frame-blocks per packet than
conversational service. This reduces the overhead from IP, UDP, and
RTP headers. However, including several frame-blocks per packet
makes the transmission more vulnerable to packet loss, so
interleaving may be used to reduce the effect of packet loss on
speech quality. A streaming server handling a large number of
clients also needs a payload format that requires as few resources as
possible when doing packetization.
For VMR-WB-enabled IP terminals at both ends, depending on the
implementation, all modes of the VMR-WB codec can be used in this
scenario. Also, both header-free and octet-aligned payload formats
(see Section 6 for details) can be utilized. For the interoperable
interconnection between VMR-WB and AMR-WB, only VMR-WB mode 3 is
used, and all restrictions described in Section 9.3 apply.
5.2. GW to IP Terminal
Another scenario occurs when VMR-WB-encoded speech will be
transmitted from a non-IP system (e.g., 3GPP2/CDMA2000 network) to an
IP terminal, and/or vice versa, as depicted in Figure 3.
VMR-WB over
3GPP2/CDMA2000 network
+------+ +----------+
| | | |
<-------------->| GW |<---------------------->| TERMINAL |
| | VMR-WB/RTP/UDP/IP | |
+------+ +----------+
|
| IP network
|
Figure 3: GW to VoIP terminal scenario
VMR-WB’s capability to switch seamlessly between operational modes is
exploited in CDMA (non-IP) networks to optimize speech quality for a
given traffic condition. To preserve this functionality in scenarios
including a gateway to an IP network using the octet-aligned payload
format, a codec mode request (CMR) field is considered. The gateway
will be responsible for forwarding the CMR between the non-IP and IP
parts in both directions. The IP terminal SHOULD follow the CMR
forwarded by the gateway to optimize speech quality going to the
non-IP decoder. The mode control algorithm in the gateway SHOULD
accommodate the delay imposed by the IP network on the response to
CMR by the IP terminal.
The IP terminal SHOULD NOT set the CMR (see Section 6.3.2), but the
gateway can set the CMR value on frames going toward the encoder in
the non-IP part to optimize speech quality from that encoder to the
gateway and to perform congestion control on the IP network.
5.3. GW to GW (between VMR-WB- and AMR-WB-Enabled Terminals)
A third likely scenario is that RTP/UDP/IP is used as transport
between two non-IP systems, i.e., IP is originated and terminated in
gateways on both sides of the IP transport, as illustrated in Figure
4. This is the most likely scenario for an interoperable
interconnection between 3GPP/(GSM-WCDMA)/AMR-WB and
3GPP2/CDMA2000/VMR-WB-enabled mobile stations. In this scenario, the
VMR-WB-enabled terminal also declares itself capable of AMR-WB with
restricted mode set as described in Section 9.3. The CMR value may be
set in packets received by the gateways on the IP network side. The
gateway should forward to the non-IP side a CMR value that is the
minimum of three values: (1) the CMR value it receives on the IP
side; (2) a CMR value it may choose for congestion control of
transmission on the IP side; and (3) the CMR value based on its
estimate of reception quality on the non-IP side. The details of the
traffic control algorithm are left to the implementation.
VMR-WB over AMR-WB over
3GPP2/CDMA2000 network 3GPP/(GSM-WCDMA) network
+------+ +------+
(AMR-WB Payload) | | AMR-WB/RTP/UDP/IP| |(AMR-WB Payload)
<---------------->| GW |<---------------->| GW |<--------------->
| | | |
+------+ +------+