Request for Comments: 3913 Microsoft
Category: Informational September 2004
Border Gateway Multicast Protocol (BGMP):
Protocol Specification
Status of this Memo
This memo provides information for the Internet community. It does
not specify an Internet standard of any kind. Distribution of this
memo is unlimited.
Copyright Notice
Copyright (C) The Internet Society (2004).
Abstract
This document describes the Border Gateway Multicast Protocol (BGMP),
a protocol for inter-domain multicast routing. BGMP builds shared
trees for active multicast groups, and optionally allows receiver
domains to build source-specific, inter-domain, distribution branches
where needed. BGMP natively supports "source-specific multicast"
(SSM). To also support "any-source multicast" (ASM), BGMP requires
that each multicast group be associated with a single root (in BGMP
it is referred to as the root domain). It requires that different
ranges of the multicast address space are associated (e.g., with
Unicast-Prefix-Based Multicast addressing) with different domains.
Each of these domains then becomes the root of the shared domain-
trees for all groups in its range. Multicast participants will
generally receive better multicast service if the session initiator’s
address allocator selects addresses from its own domain’s part of the
space, thereby causing the root domain to be local to at least one of
the session participants.
Table of Contents
1. Purpose. . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2. Terminology. . . . . . . . . . . . . . . . . . . . . . . . . . 4
3. Protocol Overview. . . . . . . . . . . . . . . . . . . . . . . 5
3.1. Design Rationale . . . . . . . . . . . . . . . . . . . . 7
4. Protocol Details . . . . . . . . . . . . . . . . . . . . . . . 8
4.1. Interaction with the EGP . . . . . . . . . . . . . . . . 8
4.2. Multicast Data Packet Processing . . . . . . . . . . . . 9
4.3. BGMP processing of Join and Prune messages and
notifications. . . . . . . . . . . . . . . . . . . . . . 10
4.3.1. Receiving Joins. . . . . . . . . . . . . . . . . 10
4.3.2. Receiving Prune Notifications. . . . . . . . . . 11
4.3.3. Receiving Route Change Notifications . . . . . . 12
4.3.4. Receiving (S,G) Poison-Reverse messages. . . . . 12
4.4. Interaction with M-IGP components. . . . . . . . . . . . 13
4.4.1. Interaction with DVMRP and PIM-DM. . . . . . . . 14
4.4.2. Interaction with PIM-SM. . . . . . . . . . . . . 15
4.4.3. Interaction with CBT . . . . . . . . . . . . . . 16
4.4.4. Interaction with MOSPF . . . . . . . . . . . . . 17
4.5. Operation over Multi-access Networks . . . . . . . . . . 17
4.6. Interaction between (S,G) state and G-routes . . . . . . 18
5. Message Formats. . . . . . . . . . . . . . . . . . . . . . . . 18
5.1. Message Header Format. . . . . . . . . . . . . . . . . . 19
5.2. OPEN Message Format. . . . . . . . . . . . . . . . . . . 19
5.3. UPDATE Message Format. . . . . . . . . . . . . . . . . . 23
5.4. Encoding examples. . . . . . . . . . . . . . . . . . . . 27
5.5. KEEPALIVE Message Format . . . . . . . . . . . . . . . . 27
5.6. NOTIFICATION Message Format. . . . . . . . . . . . . . . 28
6. BGMP Error Handling. . . . . . . . . . . . . . . . . . . . . . 30
6.1. Message Header error handling. . . . . . . . . . . . . . 30
6.2. OPEN message error handling. . . . . . . . . . . . . . . 30
6.3. UPDATE message error handling. . . . . . . . . . . . . . 31
6.4. NOTIFICATION message error handling. . . . . . . . . . . 32
6.5. Hold Timer Expired error handling. . . . . . . . . . . . 32
6.6. Finite State Machine error handling. . . . . . . . . . . 32
6.7. Cease. . . . . . . . . . . . . . . . . . . . . . . . . . 32
6.8. Connection collision detection . . . . . . . . . . . . . 32
7. BGMP Version Negotiation . . . . . . . . . . . . . . . . . . . 33
7.1. BGMP Capability Negotiation. . . . . . . . . . . . . . . 34
8. BGMP Finite State machine. . . . . . . . . . . . . . . . . . . 34
9. Security Considerations. . . . . . . . . . . . . . . . . . . . 38
10. Acknowledgements . . . . . . . . . . . . . . . . . . . . . . . 39
11. References . . . . . . . . . . . . . . . . . . . . . . . . . . 39
11.1. Normative References . . . . . . . . . . . . . . . . . . 39
11.2. Informative References . . . . . . . . . . . . . . . . . 40
Author’s Address . . . . . . . . . . . . . . . . . . . . . . . . . 40
Full Copyright Statement . . . . . . . . . . . . . . . . . . . . . 41
1. Purpose
It has been suggested that inter-domain "any-source" multicast is
better supported with a rendezvous mechanism whereby members receive
sources’ data packets without any sort of global broadcast (e.g.,
MSDP broadcasts source information, PIM-DM [PIMDM] and DVMRP [DVMRP]
broadcast initial data packets, and MOSPF [MOSPF] broadcasts
membership information). PIM-SM [PIMSM] and CBT [CBT] use a shared
group-tree, to which all members join and thereby hear from all
sources (and to which non-members do not join and thereby hear from
no sources).
This document describes BGMP, a protocol for inter-domain multicast
routing. BGMP natively supports "source-specific multicast" (SSM).
To also support "any-source multicast" (ASM), BGMP builds shared
trees for active multicast groups, and allows domains to build
source-specific, inter-domain, distribution branches where needed.
Building upon concepts from PIM-SM and CBT, BGMP requires that each
global multicast group be associated with a single root. However, in
BGMP, the root is an entire exchange or domain, rather than a single
router.
For non-source-specific groups, BGMP assumes that ranges of the
multicast address space have been associated (e.g., with Unicast-
Prefix-Based Multicast [V4PREFIX,V6PREFIX] addressing) with selected
domains. Each such domain then becomes the root of the shared
domain-trees for all groups in its range. An address allocator will
generally achieve better distribution trees if it takes its multicast
addresses from its own domain’s part of the space, thereby causing
the root domain to be local.
BGMP uses TCP as its transport protocol. This eliminates the need to
implement message fragmentation, retransmission, acknowledgement, and
sequencing. BGMP uses TCP port 264 for establishing its connections.
This port is distinct from BGP’s port to provide protocol
independence, and to facilitate distinguishing between protocol
packets (e.g., by packet classifiers, diagnostic utilities, etc.)
Two BGMP peers form a TCP connection between one another, and
exchange messages to open and confirm the connection parameters.
They then send incremental Join/Prune Updates as group memberships
change. BGMP does not require periodic refresh of individual
entries. KeepAlive messages are sent periodically to ensure the
liveness of the connection. Notification messages are sent in
response to errors or special conditions. If a connection encounters
an error condition, a notification message is sent and the connection
is closed if the error is a fatal one.
2. Terminology
This document uses the following technical terms:
Domain:
A set of one or more contiguous links and zero or more routers
surrounded by one or more multicast border routers. Note that
this loose definition of domain also applies to an external link
between two domains, as well as an exchange.
Root Domain:
When constructing a shared tree of domains for some group, one
domain will be the "root" of the tree. The root domain receives
data from each sender to the group, and functions as a rendezvous
domain toward which member domains can send inter-domain joins,
and to which sender domains can send data.
Multicast RIB:
The Routing Information Base, or routing table, used to calculate
the "next-hop" towards a particular address for multicast traffic.
Multicast IGP (M-IGP):
A generic term for any multicast routing protocol used for tree
construction within a domain. Typical examples of M-IGPs are:
PIM-SM, PIM-DM, DVMRP, MOSPF, and CBT.
EGP: A generic term for the interdomain unicast routing protocol in
use.
Typically, this will be some version of BGP which can support a
Multicast RIB, such as MBGP [MBGP], containing both unicast and
multicast address prefixes.
Component:
The portion of a border router associated with (and logically
inside) a particular domain that runs the multicast IGP (M-IGP)
for that domain, if any. Each border router thus has zero or more
components inside routing domains. In addition, each border
router with external links that do not fall inside any routing
domain will have an inter-domain component that runs BGMP.
External peer:
A border router in another multicast AS (autonomous system, as
used in BGP), to which a BGMP TCP-connection is open. If BGP is
being used as the EGP, a separate "eBGP" TCP-connection will also
be open to the same peer.
Internal peer:
Another border router of the same multicast AS. If BGP is being
used as the EGP, the border router either speaks iBGP ("internal"
BGP) directly to internal peers in a full mesh, or indirectly
through a route reflector [REFLECT].
Next-hop peer:
The next-hop peer towards a given IP address is the next EGP
router on the path to the given address, according to multicast
RIB routes in the EGP’s routing table (e.g., in MBGP, routes whose
Subsequent Address Family Identifier field indicates that the
route is valid for multicast traffic).
target:
Either an EGP peer, or an M-IGP component.
Tree State Table:
This is a table of (S-prefix,G) and (*,G-prefix) entries that have
been explicitly joined by a set of targets. Each entry has, in
addition to the source and group addresses and masks, a list of
targets that have explicitly requested data (on behalf of directly
connected hosts or downstream routers). (S,G) entries also have
an "SPT" bit.
The key words "MUST", "MUST NOT", "SHOULD", "SHOULD NOT", and "MAY"
in this document are to be interpreted as described in [RFC2119].
3. Protocol Overview
BGMP maintains group-prefix state in response to messages from BGMP
peers and notifications from M-IGP components. Group-shared trees
are rooted at the domain advertising the group prefix covering those
groups. When a receiver joins a specific group address, the border
router towards the root domain generates a group-specific Join
message, which is then forwarded Border-Router-by-Border-Router
towards the root domain (see Figure 1). BGMP Join and Prune messages
are sent over TCP connections between BGMP peers, and BGMP protocol
state is refreshed by KEEPALIVE messages periodically sent over TCP.
BGMP routers build group-specific bidirectional forwarding state as
they process the BGMP Join messages. Bidirectional forwarding state
means that packets received from any target are forwarded to all
other targets in the target list without any RPF checks. No group-
specific state or traffic exists in parts of the network where there
are no members of that group.
BGMP routers optionally build source-specific unidirectional
forwarding state, only where needed, to be compatible with source-
specific trees (SPTs) used by some M-IGPs (e.g., DVMRP, PIM-DM, or
PIM-SM), or to construct trees for source-specific groups. A domain
that uses an SPT-based M-IGP may need to inject multicast packets
from external sources via different border routers (to be compatible
with the M-IGP RPF checks) which thus act as "surrogates". For
example, in the Transit_1 domain, data from Src_A arrives at BR12,
but must be injected by BR11. A surrogate router may create a
source-specific BGMP branch if no shared tree state exists. Note:
stub domains with a single border router, such as Rcvr_Stub_7 in
Figure 1, receive all multicast data packets through that router, to
which all RPF checks point. Therefore, stub domains never build
source-specific state.
Root_Domain
[BR91]--------------------------\
| |
[BR32] [BR41]
Transit_3 Transit_4
[BR31] [BR42] [BR43]
| | |
[BR22] [BR52] [BR53]
Transit_2 Transit_5
[BR21] [BR51]
| |
[BR12] [BR61]
Transit_1[BR11]----------[BR62]Stub_6
[BR13] (Src_A)
| (Rcvr_D)
-------------------
| |
[BR71] [BR81]
Rcvr_Stub_7 Src_only_Stub_8
(Rcvr_C) (Src_B)
Figure 1: Example inter-domain topology. [BRxy] represents a BGMP
border router. Transit_X is a transit domain network. *_Stub_X is a
stub domain network.
Data packets are forwarded based on a combination of BGMP and M-IGP
rules. The router forwards to a set of targets according to a
matching (S,G) BGMP tree state entry if it exists. If not found, the
router checks for a matching (*,G) BGMP tree state entry. If neither
is found, then the packet is sent natively to the next-hop EGP peer
for G, according to the Multicast RIB (for example, in the case of a
non-member sender such as Src_B in Figure 1). If a matching entry
was found, the packet is forwarded to all other targets in the target
list. In this way BGMP trees forward data in a bidirectional manner.
If a target is an M-IGP component then forwarding is subject to the
rules of that M-IGP protocol.
3.1. Design Rationale
Several other protocols, or protocol proposals, build shared trees
within domains [PIMSM, CBT]. The design choices made for BGMP result
from our focus on Inter-Domain multicast in particular. The design
choices made by PIM-SM and CBT are better suited to the wide-area
intra-domain case. There are three major differences between BGMP
and other shared-tree protocols:
(1) Unidirectional vs. Bidirectional trees
Bidirectional trees (using bidirectional forwarding state as
described above) minimize third party dependence which is essential
in the inter-domain context. For example, in Figure 1, stub domains
7 and 8 would like to exchange multicast packets without being
dependent on the quality of connectivity of the root domain.
However, unidirectional shared trees (i.e., those using RPF checks)
have more aggressive loop prevention and share the same processing
rules as source-specific entries which are inherently unidirectional.
The lack of third party dependence concerns in the INTRA domain case
reduces the incentive to employ bidirectional trees. BGMP supports
bidirectional trees because it has to, and because it can without
excessive cost.
(2) Source-specific distribution trees/branches
In a departure from other shared tree protocols, source-specific BGMP
state is built ONLY where (a) it is needed to pull the multicast
traffic down to a BGMP router that has source-specific (S,G) state,
and (b) that router is NOT already on the shared tree (i.e., has no
(*,G) state), and (c) that router does not want to receive packets
via encapsulation from a router which is on the shared tree. BGMP
provides source-specific branches because most M-IGP protocols in use
today build source-specific trees. BGMP’s source-specific branches
eliminate the unnecessary overhead of encapsulations for high data
rate sources from the shared tree’s ingress router to the surrogate
injector (e.g., from BR12 to BR11 in Figure 1). Moreover, cases in
which shared paths are significantly longer than SPT paths will also
benefit.
However, except for source-specific group distribution trees, we do
not build source-specific inter-domain trees in general because (a)
inter-domain connectivity is generally less rich than intra-domain
connectivity, so shared distribution trees should have more
acceptable path length and traffic concentration properties in the
inter-domain context, than in the intra-domain case, and (b) by
having the shared tree state always take precedence over source-
specific tree state, we avoid ambiguities that can otherwise arise.
In summary, BGMP trees are, in a sense, a hybrid between PIM-SM and
CBT trees.
(3) Method of choosing root of group shared tree
The choice of a group’s shared-tree-root has implications for
performance and policy. In the intra-domain case it is sometimes
assumed that all potential shared-tree roots (RPs/Cores) within the
domain are equally suited to be the root for a group that is
initiated within that domain. In the INTER-domain case, there is far
more opportunity for unacceptably poor locality, and administrative
control of a group’s shared-tree root. Therefore in the intra-domain
case, other protocols sometimes treat all candidate roots (RPs or
Cores) as equivalent and emphasize load sharing and stability to
maximize performance. In the Inter-Domain case, all roots are not
equivalent, and we adopt an approach whereby a group’s root domain is
not random but is subject to administrative control.
4. Protocol Details
In this section, we describe the detailed protocol that border
routers perform. We assume that each border router conforms to the
component-based model described in [INTEROP], modulo one correction
to section 3.2 ("BGMP" Dispatcher), as follows:
The iif owner of a (*,G) entry is the component owning the next-hop
interface towards the nominal root of G, in the multicast RIB.
4.1. Interaction with the EGP
The fundamental requirements imposed by BGMP are that:
(1) For a given source-specific group and source, BGMP must be able
to look up the next-hop towards the source in the Multicast
RIB, and
(2) For a given non-source-specific group, BGMP will map the group
address to a nominal "root" address, and must be able to look
up the next-hop towards that address in the Multicast RIB.
BGMP determines the nominal "root" address as follows. If the
multicast address is a Unicast-Prefix-based Multicast address, then
the nominal root address is the embedded unicast prefix, padded with
a suffix of 0 bits to form a full address.
For example, if the IPv6 group address is
ff2e:0100:1234:5678:9abc:def0::123, then the unicast prefix is
1234:5678:9abc:def0/64, and the nominal root address would be
1234:5678:9abc:def0::. (This address is in fact the subnet router
anycast address [IPv6AA].)
Support for any-source-multicast using any address other than a
Unicast-prefix-based Multicast Address is outside the scope of this
document.
4.2. Multicast Data Packet Processing
For BGMP rules to be applied, an incoming packet must first be
"accepted":
o If the packet arrived on an interface owned by an M-IGP, the M-IGP
component determines whether the packet should be accepted or
dropped according to its rules. If the packet is accepted, the
packet is forwarded (or not forwarded) out any other interfaces
owned by the same component, as specified by the M-IGP.
o If the packet was received over a point-to-point interface owned
by BGMP, the packet is accepted.
o If the packet arrived on a multiaccess network interface owned by
BGMP, the packet is accepted if it is receiving data on a source-
specific branch, if it is the designated forwarder for the longest
matching route for S, or for the longest matching route for the
nominal root of G.
If the packet is accepted, then the router checks the tree state
table for a matching (S,G) entry. If one is found, but the packet
was not received from the next hop target towards S (if the entry’s
SPT bit is True), or was not received from the next hop target
towards G (if the entry’s SPT bit is False) then the packet is
dropped and no further actions are taken. If no (S,G) entry was
found, the router then checks for a matching (*,G) entry.
If neither is found, then the packet is forwarded towards the next-
hop peer for the nominal root of G, according to the Multicast RIB.
If a matching entry was found, the packet is forwarded to all other
targets in the target list.
Forwarding to a target which is an M-IGP component means that the
packet is forwarded out any interfaces owned by that component
according to that component’s multicast forwarding rules.
4.3. BGMP processing of Join and Prune messages and notifications
4.3.1. Receiving Joins
When the BGMP component receives a (*,G) or (S,G) Join alert from
another component, or a BGMP (S,G) or (*,G) Join message from an
external peer, it searches the tree state table for a matching entry.
If an entry is found, and that peer is already listed in the target
list, then no further actions are taken.
Otherwise, if no (*,G) or (S,G) entry was found, one is created. In