described in [RFC2212].
7. Support of the Diffserv class selectors [RFC2474] suggests that
the subnet might consider mechanisms that support priorities.
10. Fairness vs Performance
Subnetwork designers should be aware of the tradeoffs between
fairness and efficiency inherent in many transmission scheduling
algorithms. For example, many local area networks use contention
protocols to resolve access to a shared transmission channel. These
protocols represent overhead. While limiting the amount of data that
a subnet node may transmit per contention cycle helps assure timely
access to the channel for each subnet node, it also increases
contention overhead per unit of data sent.
In some mobile radio networks, capacity is limited by interference,
which in turn depends on average transmitter power. Some receivers
may require considerably more transmitter power (generating more
interference and consuming more channel capacity) than others.
In each case, the scheduling algorithm designer must balance
competing objectives: providing a fair share of capacity to each
subnet node while maximizing the total capacity of the network. One
approach for balancing performance and fairness is outlined in
[ES00].
11. Delay Characteristics
The TCP sender bases its retransmission timeout (RTO) on measurements
of the round trip delay experienced by previous packets. This allows
TCP to adapt automatically to the very wide range of delays found on
the Internet. The recommended algorithms are described in [RFC2988].
Evaluations of TCP’s retransmission timer can be found in [AP99] and
[LS00].
These algorithms model the delay along an Internet path as a
normally-distributed random variable with a slowly-varying mean and
standard deviation. TCP estimates these two parameters by
exponentially smoothing individual delay measurements, and it sets
the RTO to the estimated mean delay plus some fixed number of
standard deviations. (The algorithm actually uses mean deviation as
an approximation to standard deviation, because it is easier to
compute.)
The goal is to compute an RTO that is small enough to detect and
recover from packet losses while minimizing unnecessary ("spurious")
retransmissions when packets are unexpectedly delayed but not lost.
Although these goals conflict, the algorithm works well when the
delay variance along the Internet path is low, or the packet loss
rate is low.
If the path delay variance is high, TCP sets an RTO that is much
larger than the mean of the measured delays. If the packet loss rate
is low, the large RTO is of little consequence, as timeouts occur
only rarely. Conversely, if the path delay variance is low, then TCP
recovers quickly from lost packets; again, the algorithm works well.
However, when delay variance and the packet loss rate are both high,
these algorithms perform poorly, especially when the mean delay is
also high.
Because TCP uses returning acknowledgments as a "clock" to time the
transmission of additional data, excessively high delays (even if the
delay variance is low) also affect TCP’s ability to fully utilize a
high-speed transmission pipe. It also slows the recovery of lost
packets, even when delay variance is small.
Subnetwork designers should therefore minimize all three parameters
(delay, delay variance, and packet loss) as much as possible.
In many subnetworks, these parameters are inherently in conflict.
For example, on a mobile radio channel, the subnetwork designer can
use retransmission (ARQ) and/or forward error correction (FEC) to
trade off delay, delay variance, and packet loss in an effort to
improve TCP performance. While ARQ increases delay variance, FEC
does not. However, FEC (especially when combined with interleaving)
often increases mean delay, even on good channels where ARQ
retransmissions are not needed and ARQ would not increase either the
delay or the delay variance.
The tradeoffs among these error control mechanisms and their
interactions with TCP can be quite complex, and are the subject of
much ongoing research. We therefore recommend that subnetwork
designers provide as much flexibility as possible in the
implementation of these mechanisms, and provide access to them as
discussed above in the section on Quality of Service.
12. Bandwidth Asymmetries
Some subnetworks may provide asymmetric bandwidth (or may cause TCP
packet flows to experience asymmetry in the capacity) and the
Internet protocol suite will generally still work fine. However,
there is a case when such a scenario reduces TCP performance. Since
TCP data segments are "clocked" out by returning acknowledgments, TCP
senders are limited by the rate at which ACKs can be returned
[BPK98]. Therefore, when the ratio of the available capacity of the
Internet path carrying the data to the bandwidth of the return path
of the acknowledgments is too large, the slow return of the ACKs
directly impacts performance. Since ACKs are generally smaller than
data segments, TCP can tolerate some asymmetry, but as a general
rule, designers of subnetworks should be aware that subnetworks with
significant asymmetry can result in reduced performance, unless
issues are taken to mitigate this [RFC3449].
Several strategies have been identified for reducing the impact of
asymmetry of the network path between two TCP end hosts, e.g.,
[RFC3449]. These techniques attempt to reduce the number of ACKs
transmitted over the return path (low bandwidth channel) by changes
at the end host(s), and/or by modification of subnetwork packet
forwarding. While these solutions may mitigate the performance
issues caused by asymmetric subnetworks, they do have associated cost
and may have other implications. A fuller discussion of strategies
and their implications is provided in [RFC3449].
13. Buffering, flow and congestion control
Many subnets include multiple links with varying traffic demands and
possibly different transmission speeds. At each link there must be a
queuing system, including buffering, scheduling, and a capability to
discard excess subnet packets. These queues may also be part of a
subnet flow control or congestion control scheme.
For the purpose of this discussion, we talk about packets without
regard to whether they refer to a complete IP packet or a subnetwork
frame. At each queue, a packet experiences a delay that depends on
competing traffic and the scheduling discipline, and is subjected to
a local discarding policy.
Some subnets may have flow or congestion control mechanisms in
addition to packet dropping. Such mechanisms can operate on
components in the subnet layer, such as schedulers, shapers, or
discarders, and can affect the operation of IP forwarders at the
edges of the subnet. However, with the exception of Explicit
Congestion Notification [RFC3168] (discussed below), IP has no way to
pass explicit congestion or flow control signals to TCP.
TCP traffic, especially aggregated TCP traffic, is bursty. As a
result, instantaneous queue depths can vary dramatically, even in
nominally stable networks. For optimal performance, packets should
be dropped in a controlled fashion, not just when buffer space is
unavailable. How much buffer space should be supplied is still a
matter of debate, but as a rule of thumb, each node should have
enough buffering to hold one link_bandwidth*link_delay product’s
worth of data for each TCP connection sharing the link.
This is often difficult to estimate, since it depends on parameters
beyond the subnetwork’s control or knowledge. Internet nodes
generally do not implement admission control policies, and cannot
limit the number of TCP connections that use them. In general, it is
wise to err in favor of too much buffering rather than too little.
It may also be useful for subnets to incorporate mechanisms that
measure propagation delays to assist in buffer sizing calculations.
There is a rough consensus in the research community that active
queue management is important to improving fairness, link
utilization, and throughput [RFC2309]. Although there are questions
and concerns about the effectiveness of active queue management
(e.g., [MBDL99]), it is widely considered an improvement over tail-
drop discard policies.
One form of active queue management is the Random Early Detection
(RED) algorithm [RED93], a family of related algorithms. In one
version of RED, an exponentially-weighted moving average of the queue
depth is maintained:
When this average queue depth is between a maximum threshold
max_th and a minimum threshold min_th, the probability of packets
that are dropped is proportional to the amount by which the
average queue depth exceeds min_th.
When this average queue depth is equal to max_th, the drop
probability is equal to a configurable parameter max_p.
When this average queue depth is greater than max_th, packets are
always dropped.
Numerous variants on RED appear in the literature, and there are
other active queue management algorithms which claim various
advantages over RED [GM02].
With an active queue management algorithm, dropped packets become a
feedback signal to trigger more appropriate congestion behavior by
the TCPs in the end hosts. Randomization of dropping tends to break
up the observed tendency of TCP windows belonging to different TCP
connections to become synchronized by correlated drops, and it also
imposes a degree of fairness on those connections that implement TCP
congestion avoidance properly. Another important property of active
queue management algorithms is that they attempt to keep average
queue depths short while accommodating large short-term bursts.
Since TCP neither knows nor cares whether congestive packet loss
occurs at the IP layer or in a subnet, it may be advisable for
subnets that perform queuing and discarding to consider implementing
some form of active queue management. This is especially true if
large aggregates of TCP connections are likely to share the same
queue. However, active queue management may be less effective in the
case of many queues carrying smaller aggregates of TCP connections,
e.g., in an ATM switch that implements per-VC queuing.
Note that the performance of active queue management algorithms is
highly sensitive to settings of configurable parameters, and also to
factors such as RTT [MBB00] [FB00].
Some subnets, most notably ATM, perform segmentation and reassembly
at the subnetwork edges. Care should be taken here in designing
discard policies. If the subnet discards a fragment of an IP packet,
then the remaining fragments become an unproductive load on the
subnet that can markedly degrade end-to-end performance [RF95].
Subnetworks should therefore attempt to discard these extra fragments
whenever one of them must be discarded. If the IP packet has already
been partially forwarded when discarding becomes necessary, then
every remaining fragment except the one marking the end of the IP
packet should also be discarded. For ATM subnets, this specifically
means using Early Packet Discard and Partial Packet Discard [ATMFTM].
Some subnets include flow control mechanisms that effectively require
that the rate of traffic flows be shaped upon entry to the subnet.
One example of such a subnet mechanism is in the ATM Available Bit
rate (ABR) service category [ATMFTM]. Such flow control mechanisms
have the effect of making the subnet nearly lossless by pushing
congestion into the IP routers at the edges of the subnet. In such a
case, adequate buffering and discard policies are needed in these
routers to deal with a subnet that appears to have varying bandwidth.
Whether there is a benefit in this kind of flow control is
controversial; there are numerous simulation and analytical studies
that go both ways. It appears that some of the issues leading to
such different results include sensitivity to ABR parameters, use of
binary rather than explicit rate feedback, use (or not) of per-VC
queuing, and the specific ATM switch algorithms selected for the
study. Anecdotally, some large networks that used IP over ABR to
carry TCP traffic have claimed it to be successful, but have
published no results.
Another possible approach to flow control in the subnet would be to
work with TCP Explicit Congestion Notification (ECN) semantics
[RFC3168] through utilizing explicit congestion indicators in subnet
frames. Routers at the edges of the subnet, rather than shaping,
would set the explicit congestion bit in those IP packets that are
received in subnet frames that have an ECN indication. Nodes in the
subnet would need to implement an active queue management protocol
that marks subnet frames instead of dropping them.
ECN is currently a proposed standard, but it is not yet widely
deployed.
14. Compression
Application data compression is a function that can usually be
omitted in the subnetwork. The endpoints typically have more CPU and
memory resources to run a compression algorithm and a better
understanding of what is being compressed. End-to-end compression
benefits every network element in the path, while subnetwork-layer
compression, by definition, benefits only a single subnetwork.
Data presented to the subnetwork layer may already be in a compressed
format (e.g., a JPEG file), compressed at the application layer
(e.g., the optional "gzip", "compress", and "deflate" compression in
HTTP/1.1 [RFC2616]), or compressed at the IP layer (the IP Payload
Compression Protocol [RFC3173] supports DEFLATE [RFC2394] and LZS
[RFC2395]). Compression at the subnetwork edges is of no benefit for
any of these cases.
The subnetwork may also process data that has been encrypted by the
application (OpenPGP [RFC2440] or S/MIME [RFC2633]), just above TCP
(SSL, TLS [RFC2246]), or just above IP (IPsec ESP [RFC2406]).
Ciphers generate high-entropy bit streams lacking any patterns that
can be exploited by a compression algorithm.
However, much data is still transmitted uncompressed over the
Internet, so subnetwork compression may be beneficial. Any
subnetwork compression algorithm must not expand uncompressible data,
e.g., data that has already been compressed or encrypted.
We make a strong recommendation that subnetworks operating at low
speed or with small MTUs compress IP and transport-level headers (TCP
and UDP) using several header compression schemes developed within
the IETF [RFC3150]. An uncompressed 40-byte TCP/IP header takes
about 33 milliseconds to send at 9600 bps. "VJ" TCP/IP header
compression [RFC1144] compresses most headers to 3-5 bytes, reducing
transmission time to several milliseconds on dialup modem links.
This is especially beneficial for small, latency-sensitive packets in
interactive sessions.
Similarly, RTP compression schemes, such as CRTP [RFC2508] and ROHC
[RFC3095], compress most IP/UDP/RTP headers to 1-4 bytes. The
resulting savings are especially significant when audio packets are
kept small to minimize store-and-forward latency.
Designers should consider the effect of the subnetwork error rate on
the performance of header compression. TCP ordinarily recovers from
lost packets by retransmitting only those packets that were actually
lost; packets arriving correctly after a packet loss are kept on a
resequencing queue and do not need to be retransmitted. In VJ TCP/IP
[RFC1144] header compression, however, the receiver cannot explicitly
notify a sender of data corruption and subsequent loss of
synchronization between compressor and decompressor. It relies
instead on TCP retransmission to re-synchronize the decompressor.
After a packet is lost, the decompressor must discard every
subsequent packet, even if the subnetwork makes no further errors,
until the sending TCP retransmits to re-synchronize the decompressor.
This effect can substantially magnify the effect of subnetwork packet
losses if the sending TCP window is large, as it will often be on a
path with a large bandwidth*delay product [LRKOJ99].
Alternate header compression schemes, such as those described in
[RFC2507], include an explicit request for retransmission of an
uncompressed packet to allow decompressor resynchronization without
waiting for a TCP retransmission. However, these schemes are not yet
in widespread use.
Both TCP header compression schemes do not compress widely-used TCP
options such as selective acknowledgements (SACK). Both fail to
compress TCP traffic that makes use of explicit congestion
notification (ECN). Work is under way in the IETF ROHC WG to address
these shortcomings in a ROHC header compression scheme for TCP
[RFC3095] [RFC3096].
The subnetwork error rate also is important for RTP header
compression. CRTP uses delta encoding, so a packet loss on the link
causes uncertainty about the subsequent packets, which often must be
discarded until the decompressor has notified the compressor and the
compressor has sent re-synchronizing information. This typically
takes slightly more than the end-to-end path round-trip time. For
links that combine significant error rates with latencies that
require multiple packets to be in flight at a time, this leads to
significant error propagation, i.e., subsequent losses caused by an
initial loss.
For links that are both high-latency (multiple packets in flight from
a typical RTP stream) and error-prone, RTP ROHC provides a more
robust way of RTP header compression, at a cost of higher complexity
at the compressor and decompressor. For example, within a talk
spurt, only extended losses of (depending on the mode chosen) 12-64
packets typically cause error propagation.
15. Packet Reordering
The Internet architecture does not guarantee that packets will arrive
in the same order in which they were originally transmitted;
transport protocols like TCP must take this into account.
However, reordering does come at a cost with TCP as it is currently
defined. Because TCP returns a cumulative acknowledgment (ACK)
indicating the last in-order segment that has arrived, out-of-order
segments cause a TCP receiver to transmit a duplicate acknowledgment.
When the TCP sender notices three duplicate acknowledgments, it
assumes that a segment was dropped by the network and uses the fast
retransmit algorithm [Jac90] [RFC2581] to resend the segment. In
addition, the congestion window is reduced by half, effectively
halving TCP’s sending rate. If a subnetwork reorders segments
significantly such that three duplicate ACKs are generated, the TCP
sender needlessly reduces the congestion window and performance
suffers.
Packet reordering frequently occurs in parts of the Internet, and it
seems to be difficult or impossible to eliminate [BPS99]. For this
reason, research on improving TCP’s behavior in the face of packet
reordering [LK00] [BA02] has begun.
[BPS99] cites reasons why it may even be undesirable to eliminate
reordering. There are situations where average packet latency can be
reduced, link efficiency can be increased, and/or reliability can be
improved if reordering is permitted. Examples include certain high
speed switches within the Internet backbone and the parallel links
used over many Internet paths for load splitting and redundancy.
This suggests that subnetwork implementers should try to avoid packet
reordering whenever possible, but not if doing so compromises
efficiency, impairs reliability, or increases average packet delay.
Note that every header compression scheme currently standardized for
the Internet requires in-order packet delivery on the link between
compressor and decompressor. PPP is frequently used to carry
compressed TCP/IP packets; since it was originally designed for
point-to-point and dialup links, it is assumed to provide in-order
delivery. For this reason, subnetwork implementers who provide PPP
interfaces to VPNs and other more complex subnetworks, must also
maintain in-order delivery of PPP frames.
16. Mobility
Internet users are increasingly mobile. Not only are many Internet
nodes laptop computers, but pocket organizers and mobile embedded
systems are also becoming nodes on the Internet. These nodes may
connect to many different access points on the Internet over time,
and they expect this to be largely transparent to their activities.
Except when they are not connected to the Internet at all, and for
performance differences when they are connected, they expect that
everything will "just work" regardless of their current Internet
attachment point or local subnetwork technology.
Changing a host’s Internet attachment point involves one or more of
the following steps.
First, if use of the local subnetwork is restricted, the user’s
credentials must be verified and access granted. There are many ways
to do this. A trivial example would be an "Internet cafe" that
grants physical access to the subnetwork for a fee. Subnetworks may
implement technical access controls of their own; one example is IEEE
802.11 Wireless Equivalent Privacy [IEEE80211]. It is common
practice for both cellular telephone and Internet service providers
(ISPs) to agree to serve one anothers’ users; RADIUS [RFC2865] is the
standard method for ISPs to exchange authorization information.
Second, the host may have to be reconfigured with IP parameters
appropriate for the local subnetwork. This usually includes setting
an IP address, default router, and domain name system (DNS) servers.
On multiple-access networks, the Dynamic Host Configuration Protocol
(DHCP) [RFC2131] is almost universally used for this purpose. On PPP
links, these functions are performed by the IP Control Protocol
(IPCP) [RFC1332].
Third, traffic destined for the mobile host must be routed to its
current location. This roaming function is the most common meaning
of the term "Internet mobility".
Internet mobility can be provided at any of several layers in the
Internet protocol stack, and there is ongoing debate as to which is
the most appropriate and efficient. Mobility is already a feature of
certain application layer protocols; the Post Office Protocol (POP)
[RFC1939] and the Internet Message Access Protocol (IMAP) [RFC3501]
were created specifically to provide mobility in the receipt of
electronic mail.
Mobility can also be provided at the IP layer [RFC3344]. This
mechanism provides greater transparency, viz., IP addresses that
remain fixed as the nodes move, but at the cost of potentially
significant network overhead and increased delay because of the sub-
optimal network routing and tunneling involved.
Some subnetworks may provide internal mobility, transparent to IP, as
a feature of their own internal routing mechanisms. To the extent
that these simplify routing at the IP layer, reduce the need for
mechanisms like Mobile IP, or exploit mechanisms unique to the
subnetwork, this is generally desirable. This is especially true
when the subnetwork covers a relatively small geographic area and the
users move rapidly between the attachment points within that area.
Examples of internal mobility schemes include Ethernet switching and
intra-system handoff in cellular telephony.
However, if the subnetwork is physically large and connects to other
parts of the Internet at multiple geographic points, care should be
taken to optimize the wide-area routing of packets between nodes on
the external Internet and nodes on the subnet. This is generally
done with "nearest exit" routing strategies. Because a given
subnetwork may be unaware of the actual physical location of a
destination on another subnetwork, it simply routes packets bound for
the other subnetwork to the nearest router between the two. This
implies some awareness of IP addressing and routing within the
subnetwork. The subnetwork may wish to use IP routing internally for
wide area routing and restrict subnetwork-specific routing to
constrained geographic areas where the effects of suboptimal routing
are minimized.
17. Routing
Subnetworks connecting more than two systems must provide their own
internal Layer-2 forwarding mechanisms, either implicitly (e.g.,
broadcast) or explicitly (e.g., switched). Since routing is the
major function of the Internet layer, the question naturally arises
as to the interaction between routing at the Internet layer and
routing in the subnet, and proper division of function between the
two.
Layer-2 subnetworks can be point-to-point, connecting two systems, or
multipoint. Multipoint subnetworks can be broadcast (e.g., shared
media or emulated) or non-broadcast. Generally, IP considers
multipoint subnetworks as broadcast, with shared-medium Ethernet as
the canonical (and historical) example, and point-to-point
subnetworks as a degenerate case. Non-broadcast subnetworks may
require additional mechanisms, e.g., above IP at the routing layer
[RFC2328].