Request for Comments: 3714 J. Kempf, Ed.
Category: Informational March 2004
IAB Concerns Regarding Congestion Control for
Voice Traffic in the Internet
Status of this Memo
This memo provides information for the Internet community. It does
not specify an Internet standard of any kind. Distribution of this
memo is unlimited.
Copyright Notice
Copyright (C) The Internet Society (2004). All Rights Reserved.
Abstract
This document discusses IAB concerns about effective end-to-end
congestion control for best-effort voice traffic in the Internet.
These concerns have to do with fairness, user quality, and with the
dangers of congestion collapse. The concerns are particularly
relevant in light of the absence of a widespread Quality of Service
(QoS) deployment in the Internet, and the likelihood that this
situation will not change much in the near term. This document is
not making any recommendations about deployment paths for Voice over
Internet Protocol (VoIP) in terms of QoS support, and is not claiming
that best-effort service can be relied upon to give acceptable
performance for VoIP. We are merely observing that voice traffic is
occasionally deployed as best-effort traffic over some links in the
Internet, that we expect this occasional deployment to continue, and
that we have concerns about the lack of effective end-to-end
congestion control for this best-effort voice traffic.
Table of Contents
1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . . 2
2. An Example of the Potential for Trouble. . . . . . . . . . . . 4
3. Why are Persistent, High Drop Rates a Problem? . . . . . . . . 6
3.1. Congestion Collapse. . . . . . . . . . . . . . . . . . . 6
3.2. User Quality . . . . . . . . . . . . . . . . . . . . . . 7
3.3. The Amorphous Problem of Fairness. . . . . . . . . . . . 8
4. Current efforts in the IETF. . . . . . . . . . . . . . . . . . 10
4.1. RTP. . . . . . . . . . . . . . . . . . . . . . . . . . . 10
4.2. TFRC . . . . . . . . . . . . . . . . . . . . . . . . . . 11
4.3. DCCP . . . . . . . . . . . . . . . . . . . . . . . . . . 12
4.4. Adaptive Rate Audio Codecs . . . . . . . . . . . . . . . 12
4.5. Differentiated Services and Related Topics . . . . . . . 13
5. Assessing Minimum Acceptable Sending Rates . . . . . . . . . . 13
5.1. Drop Rates at 4.75 kbps Minimum Sending Rate . . . . . . 17
5.2. Drop Rates at 64 kbps Minimum Sending Rate . . . . . . . 18
5.3. Open Issues. . . . . . . . . . . . . . . . . . . . . . . 18
5.4. A Simple Heuristic . . . . . . . . . . . . . . . . . . . 19
6. Constraints on VoIP Systems . . . . . . . . . . . . . . . . . . 20
7. Conclusions and Recommendations. . . . . . . . . . . . . . . . 20
8. Acknowledgements . . . . . . . . . . . . . . . . . . . . . . . 21
9. References . . . . . . . . . . . . . . . . . . . . . . . . . . 21
9.1. Normative References . . . . . . . . . . . . . . . . . . 21
9.2. Informative References . . . . . . . . . . . . . . . . . 22
10. Appendix - Sending Rates with Packet Drops . . . . . . . . . . 26
11. Security Considerations. . . . . . . . . . . . . . . . . . . . 29
12. IANA Considerations. . . . . . . . . . . . . . . . . . . . . . 29
13. Authors’ Addresses . . . . . . . . . . . . . . . . . . . . . . 30
14. Full Copyright Statement . . . . . . . . . . . . . . . . . . . 31
1. Introduction
While many in the telephony community assume that commercial VoIP
service in the Internet awaits effective end-to-end QoS, in reality
voice service over best-effort broadband Internet connections is an
available service now with growing demand. While some ISPs deploy
QoS on their backbones, and some corporate intranets offer end-to-end
QoS internally, end-to-end QoS is not generally available to
customers in the current Internet. Given the current commercial
interest in VoIP on best-effort media connections, it seems prudent
to examine the potential effect of real time flows on congestion. In
this document, we perform such an analysis. Note, however, that this
document is not making any recommendations about deployment paths for
VoIP in terms of QoS support, and is not claiming that best-effort
service can be relied upon to give acceptable performance for VoIP.
This document is also not discussing signalling connections for VoIP.
However, voice traffic is in fact occasionally deployed as best
effort traffic over some links in the Internet today, and we expect
this occasional deployment to continue. This document expresses our
concern over the lack of effective end-to-end congestion control for
this best-effort voice traffic.
Assuming that VoIP over best-effort Internet connections continues to
gain popularity among consumers with broadband connections, the
deployment of end-to-end QoS mechanisms in public ISPs may be slow.
The IETF has developed standards for QoS mechanisms in the Internet
[DIFFSERV, RSVP] and continues to be active in this area [NSIS,COPS].
However, the deployment of technologies requiring change to the
Internet infrastructure is subject to a wide range of commercial as
well as technical considerations, and technologies that can be
deployed without changes to the infrastructure enjoy considerable
advantages in the speed of deployment. RFC 2990 outlines some of the
technical challenges to the deployment of QoS architectures in the
Internet [RFC2990]. Often, interim measures that provide support for
fast-growing applications are adopted, and are successful enough at
meeting the need that the pressure for a ubiquitous deployment of the
more disruptive technologies is reduced. There are many examples of
the slow deployment of infrastructure that are similar to the slow
deployment of QoS mechanisms, including IPv6, IP multicast, or of a
global PKI for IKE and IPsec support.
Interim QoS measures that can be deployed most easily include
single-hop or edge-only QoS mechanisms for VoIP traffic on individual
congested links, such as edge-only QoS mechanisms for cable access
networks. Such local forms of QoS could be quite successful in
protecting some fraction of best-effort VoIP traffic from congestion.
However, these local forms of QoS are not directly visible to the
end-to-end VoIP connection. A best-effort VoIP connection could
experience high end-to-end packet drop rates, and be competing with
other best-effort traffic, even if some of the links along the path
might have single-hop QoS mechanisms.
The deployment of IP telephony is likely to include best-effort
broadband connections to public-access networks, in addition to other
deployment scenarios of dedicated IP networks, or as an alternative
to band splitting on the last mile of ADSL deployments or QoS
mechanisms on cable access networks. There already exists a
rapidly-expanding deployment of VoIP services intended to operate
over residential broadband access links (e.g., [FWD, Vonage]). At
the moment, many public-access IP networks are uncongested in the
core, with low or moderate levels of link utilization, but this is
not necessarily the case on last hop links. If an IP telephony call
runs completely over the Internet, the connection could easily
traverse congested links on both ends. Because of economic factors,
the growth rate of Internet telephony is likely to be greatest in
developing countries, where core links are more likely to be
congested, making congestion control an especially important topic
for developing countries.
Given the possible deployment of IP telephony over congested best-
effort networks, some concerns arise about the possibilities of
congestion collapse due to a rapid growth in real-time voice traffic
that does not practice end-to-end congestion control. This document
raises some concerns about fairness, user quality, and the danger of
congestion collapse that would arise from a rapid growth in best-
effort telephony traffic on best-effort networks. We consider best-
effort telephony connections that have a minimum sending rate and
that compete directly with other best-effort traffic on a path with
at least one congested link, and address the specific question of
whether such traffic should be required to terminate, or to suspend
sending temporarily, in the face of a persistent, high packet drop
rate, when reducing the sending rate is not a viable alternative.
The concerns in this document about fairness and the danger of
congestion collapse apply not only to telephony traffic, but also to
video traffic and other best-effort real-time traffic with a minimum
sending rate. RFC 2914 already makes the point that best-effort
traffic requires end-to-end congestion control [RFC2914]. Because
audio traffic sends at such a low rate, relative to video and other
real-time traffic, it is sometimes claimed that audio traffic doesn’t
require end-to-end congestion control. Thus, while the concerns in
this document are general, the document focuses on the particular
issue of best-effort audio traffic.
Feedback can be sent to the IAB mailing list at iab@ietf.org, or to
the editors at floyd@icir.org and kempf@docomolabs-usa.com. Feedback
can also be sent to the end2end-interest mailing list [E2E].
2. An Example of the Potential for Trouble
At the November, 2002, IEPREP Working Group meeting in Atlanta, a
brief demonstration was made of VoIP over a shared link between a
hotel room in Atlanta, Georgia, USA, and Nairobi, Kenya. The link
ran over the typical uncongested Internet backbone and access links
to peering points between either endpoint and the Internet backbone.
The voice quality on the call was very good, especially in comparison
to the typical quality obtained by a circuit-switched call with
Nairobi. A presentation that accompanied the demonstration described
the access links (e.g., DSL, T1, T3, dialup, and cable modem links)
as the primary source of network congestion, and described VoIP
traffic as being a very small percentage of the packets in commercial
ISP traffic [A02]. The presentation further stated that VoIP
received good quality in the presence of packet drop rates of 5-40%
[AUT]. The VoIP call used an ITU-T G.711 codec, plus proprietary FEC
encoding, plus RTP/UDP/IP framing. The resulting traffic load over
the Internet was substantially more than the 64 kbps required by the
codec. The primary congestion point along the path of the
demonstration was a 128 kbps access link between an ISP in Kenya and
several of its subscribers in Nairobi. So the single VoIP call
consumed more than half of the access link capacity, capacity that is
shared across several different users.
Note that this network configuration is not a particularly good one
for VoIP. In particular, if there are data services running TCP on
the link with a typical packet size of 1500 bytes, then some voice
packets could be delayed an additional 90 ms, which might cause an
increase in the end to end delay above the ITU-recommended time of
150 ms [G.114] for speech traffic. This would result in a delay
noticeable to users, with an increased variation in delay, and
therefore in call quality, as the bursty TCP traffic comes and goes.
For a call that already had high delay, such as the Nairobi call from
the previous paragraph, the increased jitter due to competing TCP
traffic also increases the requirements on the jitter buffer at the
receiver. Nevertheless, VoIP usage over congested best-effort links
is likely to increase in the near future, regardless of VoIP’s
superior performance with "carrier class" service. A best-effort
VoIP connection that persists in sending packets at 64 Kbps,
consuming half of a 128 Kbps access link, in the face of a drop rate
of 40%, with the resulting user-perceptible degradation in voice
quality, is not behaving in a way that serves the interests of either
the VoIP users or the other concurrent users of the network.
As the Nairobi connection demonstrates, prescribing universal
overprovisioning (or more precisely, provisioning sufficient to avoid
persistent congestion) as the solution to the problem is not an
acceptable generic solution. For example, in regions of the world
where circuit-switched telephone service is poor and expensive, and
Internet access is possible and lower cost, provisioning all Internet
links to avoid congestion is likely to be impractical or impossible.
In particular, an over-provisioned core is not by itself sufficient
to avoid congestion collapse all the way along the path, because an
over-provisioned core can not address the common problem of
congestion on the access links. Many access links routinely suffer
from congestion. It is important to avoid congestion collapse along
the entire end-to-end path, including along the access links (where
congestion collapse would consist of congested access links wasting
scarce bandwidth carrying packets that will only be dropped
downstream). So an over-provisioned core does not by itself
eliminate or reduce the need for end-to-end congestion avoidance and
control.
There are two possible mechanisms for avoiding this congestion
collapse: call rejection during busy periods, or the use of end-to-
end congestion control. Because there are currently no
acceptance/rejection mechanisms for best-effort traffic in the
Internet, the only alternative is the use of end-to-end congestion
control. This is important even if end-to-end congestion control is
invoked only in those very rare scenarios with congestion in
generally-uncongested access links or networks. There will always be
occasional periods of high demand, e.g., in the two hours after an
earthquake or other disaster, and this is exactly when it is
important to avoid congestion collapse.
Best-effort traffic in the Internet does not include mechanisms for
call acceptance or rejection. Instead, a best-effort network itself
is largely neutral in terms of resource management, and the
interaction of the applications’ transport sessions mutually
regulates network resources in a reasonably fair fashion. One way to
bring voice into the best-effort environment in a non-disruptive
manner is to focus on the codec and look at rate adaptation measures
that can successfully interoperate with existing transport protocols
(e.g., TCP), while at the same time preserving the integrity of a
real-time, analog voice signal; another way is to consider codecs
with fixed sending rates. Whether the codec has a fixed or variable
sending rate, we consider the appropriate response when the codec is
at its minimum data rate, and the packet drop rate experienced by the
flow remains high. This is the key issue addressed in this document.
3. Why are Persistent, High Drop Rates a Problem?
Persistent, high packet drop rates are rarely seen in the Internet
today, in the absence of routing failures or other major disruptions.
This happy situation is due primarily to low levels of link
utilization in the core, with congestion typically found on lower-
capacity access links, and to the use of end-to-end congestion
control in TCP. Most of the traffic on the Internet today uses TCP,
and TCP self-corrects so that the two ends of a connection reduce the
rate of packet sending if congestion is detected. In the sections
below, we discuss some of the problems caused by persistent, high
packet drop rates.
3.1. Congestion Collapse
One possible problem caused by persistent, high packet drop rates is
that of congestion collapse. Congestion collapse was first observed
during the early growth phase of the Internet of the mid 1980s
[RFC896], and the fix was provided by Van Jacobson, who developed the
congestion control mechanisms that are now required in TCP
implementations [Jacobson88, RFC2581].
As described in RFC 2914, congestion collapse occurs in networks with
flows that traverse multiple congested links having persistent, high
packet drop rates [RFC2914]. In particular, in this scenario packets
that are injected onto congested links squander scarce bandwidth
since these packets are only dropped later, on a downstream congested
link. If congestion collapse occurs, all traffic slows to a crawl
and nobody gets acceptable packet delivery or acceptable performance.
Because congestion collapse of this form can occur only for flows
that traverse multiple congested links, congestion collapse is a
potential problem in VoIP networks when both ends of the VoIP call
are on an congested broadband connection such as DSL, or when the
call traverses a congested backbone or transoceanic link.
3.2. User Quality
A second problem with persistent, high packet drop rates concerns
service quality seen by end users. Consider a network scenario where
each flow traverses only one congested link, as could have been the
case in the Nairobi demonstration above. For example, imagine N VoIP
flows sharing a 128 Kbps link, with each flow sending at least 64
Kbps. For simplicity, suppose the 128 Kbps link is the only
congested link, and there is no traffic on that link other than the N
VoIP calls. We will also ignore for now the extra bandwidth used by
the telephony traffic for FEC and packet headers, or the reduced
bandwidth (often estimated as 70%) due to silence suppression. We
also ignore the fact that the two streams composing a bidirectional
VoIP call, one for each direction, can in practice add to the load on
some links of the path. Given these simplified assumptions, the
arrival rate to that link is at least N*64 Kbps. The traffic
actually forwarded is at most 2*64 Kbps (the link bandwidth), so at
least (N-2)*64 Kbps of the arriving traffic must be dropped. Thus, a
fraction of at least (N-2)/N of the arriving traffic is dropped, and
each flow receives on average a fraction 1/N of the link bandwidth.
An important point to note is that the drops occur randomly, so that
no one flow can be expected statistically to present better quality
service to users than any other. Everybody’s voice quality therefore
suffers.
It seems clear from this simple example that the quality of best-
effort VoIP traffic over congested links can be improved if each VoIP
flow uses end-to-end congestion control, and has a codec that can
adapt the bit rate to the bandwidth actually received by that flow.
The overall effect of these measures is to reduce the aggregate
packet drop rate, thus improving voice quality for all VoIP users on
the link. Today, applications and popular codecs for Internet
telephony attempt to compensate by using more FEC, but controlling
the packet flow rate directly should result in less redundant FEC
information, and thus less bandwidth, thereby improving throughput
even further. The effect of delay and packet loss on VoIP in the
presence of FEC has been investigated in detail in the literature
[JS00, JS02, JS03, MTK03]. One rule of thumb is that when the packet
loss rate exceeds 20%, the audio quality of VoIP is degraded beyond
usefulness, in part due to the bursty nature of the losses [S03]. We
are not aware of measurement studies of whether VoIP users in
practice tend to hang up when packet loss rates exceed some limit.
The simple example in this section considered only voice flows, but
in reality, VoIP traffic will compete with other flows, most likely
TCP. The response of VoIP traffic to congestion works best by taking
into account the congestion control response of TCP, as is discussed
in the next subsection.
3.3. The Amorphous Problem of Fairness
A third problem with persistent, high packet drop rates is fairness.
In this document we consider fairness with regard to best-effort VoIP
traffic competing with other best-effort traffic in the Internet.
That is, we are explicitly not addressing the issues raised by
emergency services, or by QoS-enabled traffic that is known to be
treated separately from best-effort traffic at a congested link.
While fairness is a bit difficult to quantify, we can illustrate the
effect by adding TCP traffic to the congested link discussed in the
previous section. In this case, the non-congestion-controlled
traffic and congestion-controlled TCP traffic [RFC2914] share the
link, with the congestion-controlled traffic’s sending rate
determined by the packet drop rate experienced by those flows. As in
the previous section, the 128 Kbps link has N VoIP connections each
sending 64 Kbps, resulting in packet drop rate of at least (N-2)/N on
the congested link. Competing TCP flows will experience the same
packet drop rates. However, a TCP flow experiencing the same packet
drop rates will be sending considerably less than 64 Kbps. From the
point of view of who gets what amount of bandwidth, the VoIP traffic
is crowding out the TCP traffic.
Of course, this is only one way to look at fairness. The relative
fairness between VoIP and TCP traffic can be viewed several different
ways, depending on the assumptions that one makes on packet sizes and
round-trip times. In the presence of a fixed packet drop rate, for
example, a TCP flow with larger packets sends more (in Bps, bytes per
second) than a TCP flow with smaller packets, and a TCP flow with a
shorter round-trip time sends more (in Bps) than a TCP flow with a
larger round-trip time. In environments with high packet drop rates,
TCP’s sending rate depends on the algorithm for setting the
retransmit timer (RTO) as well, with a TCP implementation having a
more aggressive RTO setting sending more than a TCP implementation
having a less aggressive RTO setting.
Unfortunately, there is no obvious canonical round-trip time for
judging relative fairness of flows in the network. Agreement in the
literature is that the majority of packets on most links in the
network experience round-trip times between 10 and 500 ms [RTTWeb].
(This does not include satellite links.) As a result, if there was a
canonical round-trip for judging relative fairness, it would have to
be within that range. In the absence of a single representative
round-trip time, the assumption of this paper is that it is
reasonable to consider fairness between a VoIP connection and a TCP
connection with the same round-trip time.
Similarly, there is no canonical packet size for judging relative
fairness between TCP connections. However, because the most common
packet size for TCP data packets is 1460 bytes [Measurement], we
assume that it is reasonable to consider fairness between a VoIP
connection, and a TCP connection sending 1460-byte data packets.
Note that 1460 bytes is considerably larger than is typically used
for VoIP packets.
In the same way, while RFC 2988 specifies TCP’s algorithm for setting
TCP’s RTO, there is no canonical value for the minimum RTO, and the
minimum RTO heavily affects TCP’s sending rate in times of high
congestion [RFC2988]. RFC 2988 specifies that TCP’s RTO must be set
to SRTT + 4*RTTVAR, for SRTT the smoothed round-trip time, and for
RTTVAR the mean deviation of recent round-trip time measurements.
RFC 2988 further states that the RTO "SHOULD" have a minimum value of
1 second. However, it is not uncommon in practice for TCP
implementations to have a minimum RTO as low as 100 ms. For the
purposes of this document, in considering relative fairness, we will
assume a minimum RTO of 100 ms.
As an additional complication, TCP connections that use fine-grained