Request for Comments: 4392 IBM
Category: Informational April 2006
IP over InfiniBand (IPoIB) Architecture
Status of This Memo
This memo provides information for the Internet community. It does
not specify an Internet standard of any kind. Distribution of this
memo is unlimited.
Copyright Notice
Copyright (C) The Internet Society (2006).
Abstract
InfiniBand is a high-speed, channel-based interconnect between
systems and devices.
This document presents an overview of the InfiniBand architecture.
It further describes the requirements and guidelines for the
transmission of IP over InfiniBand. Discussions in this document are
applicable to both IPv4 and IPv6 unless explicitly specified. The
encapsulation of IP over InfiniBand and the mechanism for IP address
resolution on IB fabrics are covered in other documents.
Table of Contents
1. Introduction to InfiniBand ......................................2
1.1. InfiniBand Architecture Specification ......................2
1.2. Overview of InfiniBand Architecture ........................2
1.2.1. InfiniBand Addresses ................................6
1.2.1.1. Unicast GIDs ...............................7
1.2.1.2. Multicast GIDs .............................7
1.3. InfiniBand Multicast Group Management ......................9
1.3.1. Multicast Member Record ............................10
1.3.1.1. JoinState .................................10
1.3.2. Join and Leave Operations ..........................11
1.3.2.1. Creating a Multicast Group ................11
1.3.2.2. Deleting a Multicast Group ................11
1.3.2.3. Multicast Group Create/Delete Traps .......12
2. Management of InfiniBand Subnet ................................12
3. IP over IB .....................................................12
3.1. InfiniBand as Datalink ....................................13
3.2. Multicast Support .........................................13
3.2.1. Mapping IP Multicast to IB Multicast ...............14
3.2.2. Transient Flag in IB MGIDs .........................14
3.3. IP Subnets Across IB Subnets ..............................14
4. IP Subnets in InfiniBand Fabrics ...............................14
4.1. IPoIB VLANs ...............................................16
4.2. Multicast in IPoIB subnets ................................16
4.2.1. Sending IP Multicast Datagrams .....................17
4.2.2. Receiving Multicast Packets ........................18
4.2.3. Router Considerations for IPoIB ....................18
4.2.4. Impact of InfiniBand Architecture Limits ...........19
4.2.5. Leaving/Deleting a Multicast Group .................19
4.3. Transmission of IPoIB Packets .............................20
4.4. Reverse Address Resolution Protocol (RARP) and
Static ARP Entries ........................................20
4.5. DHCPv4 and IPoIB ..........................................21
5. QoS and Related Issues .........................................21
6. Security Considerations ........................................21
7. Acknowledgements ...............................................21
8. References .....................................................21
8.1. Normative References ......................................21
8.2. Informative References ....................................22
1. Introduction to InfiniBand
The InfiniBand Trade Association (IBTA) was formed to develop an I/O
specification to deliver a channel based, switched fabric technology.
The InfiniBand standard is aimed at meeting the requirements of
scalability, reliability, availability, and performance of servers in
data centers.
1.1. InfiniBand Architecture Specification
The InfiniBand Trade Association specification is available for
download from http://www.infinibandta.org.
1.2. Overview of InfiniBand Architecture
For a more complete overview, the reader is referred to chapter 3 of
the InfiniBand specification.
InfiniBand Architecture (IBA) defines a System Area Network (SAN) for
connecting multiple independent processor platforms, I/O platforms,
and I/O devices. The IBA SAN is a communications and management
infrastructure supporting both I/O and inter-processor communications
for one or more computer systems.
An IBA SAN consists of processor nodes and I/O units connected
through an IBA fabric made up of cascaded switches and IB routers
(connecting IB subnets). I/O units can range in complexity from a
single Application-specific Integrated Circuit (ASIC) IBA-attached
device (such as a LAN adapter) to a large, memory-rich Redundant
Array of Independent Disks (RAID) subsystem.
An IBA network may be subdivided into subnets interconnected by
routers. These are IB routers and IB subnets and not IP routers or
IP subnets. This document will refer to InfiniBand routers and
subnets as ’IB routers’ and ’IB subnets’ respectively. The IP
routers and IP subnets will be referred to as ’routers’ and
’subnets’, respectively.
Each IB node or switch may attach to a single or multiple switches or
directly with each other. Each IB unit interfaces with the link by
way of channel adapters (CAs). The architecture supports multiple
CAs per unit with each CA providing one or more ports that connect to
the fabric. Each CA appears as a node to the fabric.
The ports are the endpoints to which the data is sent. However, each
of the ports may include multiple QPs (Queue Pairs) that may be
directly addressed from a remote peer. From the point of view of
data transfer the QP number (QPN) is part of the address.
IBA supports both connection-oriented and datagram service between
the ports. The peers are identified by QPN and the port identifier.
There are a two exceptions. QPNs are not used when packets are
multicast. QPNs are also not used in the Raw Datagram mode.
A port, in a data packet, is identified by a Local Identifier (LID)
and optionally a Global Identifier (GID). The GID in the packet is
needed only when communicating across an IB subnet, though it may
always be included.
The GID is 128 bits long and is formed by the concatenation of a 64-
bit IB subnet prefix and a 64-bit EUI-64-compliant portion. The
EUI-64 portion of a GID is referred to as the Global Unique
Identifier (GUID; EUI stands for Extended Unique Identifier). The
LID is a 16-bit value that is assigned when the port becomes active.
The GUID is the only persistent identifier of a port. However, it
cannot be used as an address in a packet. If the prefix is modified,
then the GID may change. The subnet manager may attempt to keep the
LID values constant across reboots, but that is not a requirement.
The assignment of the GID and the LID is done by the subnet manager.
Every IB subnet has at least one subnet manager component that
controls the fabric. It assigns the LIDs and GIDs. The subnet
manager also programs the switches so that they route packets between
destinations. The subnet manager (SM) and a related component, the
subnet administrator (SA), are the central repository of all
information that is required to set-up and bring up the fabric.
IB routers are components that route packets between IB subnets based
on the GIDs. Thus, within an IB subnet a packet may or may not
include a GID but when going across an IB subnet the GID must be
included. A LID is always needed in a packet since the destination
within a subnet is determined by it.
A CA and a switch may have multiple ports. Each CA port is assigned
its own LID or a range of LIDs. The ports of a switch are not
addressable by LIDs/GIDs or, in other words, are transparent to other
end nodes. Each port has its own set of buffers. The buffering is
channeled through virtual lanes (VL) where each VL has its own flow
control. There may be up to 16 VLs.
VLs provide a mechanism for creating multiple virtual links within a
single physical link. All ports must support VL15 which is reserved
exclusively for subnet management datagrams and hence does not
concern the IP over Infiniband (IPoIB) discussions. The actual VL
that a packet uses is configured by the SM in the switch/channel
adapter tables and is determined based on the Service Level (SL)
specified in every packet. There are 16 possible SLs.
In addition to the features described above viz. QPs, SLs, and
addressing (GID/LID), IBA also defines the following:
Partitioning:
Every packet, but for the raw datagrams, carries the partition key
(P_Key). These values are used for isolation in the fabric. A
switch (this is an optional feature) may be programmed by the SM
to drop packets not having a certain key. The CA ports always
check for the P_Keys. A CA port may belong to multiple
partitions. P_Key checking is optional at IB routers.
A P_Key may be described as having ’limited membership’ or ’full
membership’. For a packet to be accepted, at least one of the
P_Keys (i.e., the P_Key in the packet or the P_Key in the port)
must be ’full membership’ P_Keys.
Q_Keys:
Q_Keys are used to enforce access rights for reliable and
unreliable IB datagram services. Raw datagram services do not use
Q_Keys. At communication establishment, the endpoints exchange
the Q_Keys and must always use the relevant Q_Keys when
communicating with one another. Multicast packets use the Q_Key
associated with the multicast group.
Q_Keys with the most significant bit set are considered controlled
Q_Keys (such as the General Service Interface (GSI) Q_Key
[IB_ARCH]) and a Host Channel Adapter (HCA) does not allow a
consumer to arbitrarily specify a controlled Q_Key. An attempt to
send a controlled Q_Key results in using the Q_Key in the QP
context. Thus, the Operating System maintains control since it
can configure the QP context for the controlled Q_Key for
privileged consumers. It must be noted that though the notion of
a ’controlled Q_Key’ is suggested by IB specification, it does not
require its use or implementation.
Multicast support:
A switch may support multicasting, that is, replication of packets
across multiple output ports. This is an optional feature.
Similarly, support for sending/receiving multicast packets is
optional in CAs. A multicast group is identified by a GID. The
GID format is as defined in RFC 2373 on IPv6 addressing [IB_ARCH].
Thus, from an IPv6-over-InfiniBand point of view, the data link
multicast address looks like the network address. An IB port must
explicitly join a multicast group by sending a request to the SM
to receive multicast packets. A port may send packets to any
multicast group. In both cases, the multicast LID to be used in
the packets is received from the SM.
There are six methods for data transfer in IB architecture:
1. Unreliable Datagram (unacknowledged - connectionless)
The Unreliable Datagram (UD) service is connectionless and
unacknowledged. It allows the QP to communicate with any
unreliable datagram QP on any node.
The switches and hence each link can support only a certain
MTU. The MTU ranges are 256 octets, 512 octets, 1024 octets,
2048 octets, and 4096 octets. A UD packet cannot be larger
than the link MTU between the two peers.
2. Reliable Datagram (acknowledged - multiplexed)
The Reliable Datagram (RD) service is multiplexed over
connections between nodes called End-to-End Contexts (EEC),
which allows each RD QP to communicate with any RD QP on any
node with an established EEC. Multiple QPs can use the same
EEC and a single QP can use multiple EECs (one for each remote
node per reliable datagram domain).
3. Reliable Connected (acknowledged - connection oriented)
The Reliable Connected (RC) service associates a local QP with
one and only one remote QP. The message sizes maybe as large
as 2^31 octets in length. The CA implementation takes care of
segmentation and assembly.
4. Unreliable Connected (unacknowledged - connection oriented)
The Unreliable Connected (UC) service associates one local QP
with one and only one remote QP. There is no acknowledgement
and hence no resend of lost or corrupted packets. Such packets
are therefore simply dropped. It is similar to RC otherwise.
5. Raw Ethertype (unacknowledged - connectionless)
The Ethertype raw datagram packet contains a generic transport
header that is not interpreted by the CA but it specifies the
protocol type. The values for ethertype are the same as
defined by Internet Assigned Numbers Authority (IANA) [IANA]
for ethertype.
6. Raw IPv6 (unacknowledged - connectionless)
Using IPv6 raw datagram service, the IBA CA can support
standard protocol layers atop IPv6 (such as TCP/UDP). Thus,
native IPv6 packets can be bridged into the IBA SAN and
delivered directly to a port and to its IPv6 raw datagram QP.
The first four types are referred to as IB transports. The latter
two are classified as raw datagrams. There is no indication of the
QP number in the raw datagram packets. The raw datagram packets are
limited by the link MTU in size.
The two connected modes and the Reliable Datagram mode may also
support Automatic Path Migration (APM). This is an optional facility
that provides for a hardware based path fail over. An alternate path
is associated with the QP when the connection/EE context is first
created. If unrecoverable errors are encountered, the connection
switches to using the alternative path.
1.2.1. InfiniBand Addresses
The InfiniBand architecture borrows heavily from the IPv6
architecture in terms of the InfiniBand subnet structure and GIDs.
The InfiniBand architecture defines the GID associated with a port as
a 128-bit unicast or multicast identifier. IBA derives the GID
address format, as defined in RFC 2373 [IB_ARCH], with some
additional properties/restrictions defined to facilitate efficient
discovery, communication, and routing.
Note: The IBA explicitly refers to RFC 2373, which is obsolete
[RFC3513]. It must be noted that IBA is therefore unaffected by
any further changes that are introduced in IPv6 addressing
architecture.
IBA defines two types of GIDs: unicast and multicast.
1.2.1.1. Unicast GIDs
The unicast GIDs are defined, as in IPv6, with three scopes. The IB
specification states the following:
a. link local: FE80/10.
The IB routers will not forward packets with a link-
local address in source or destination beyond the IB
subnet.
b. site local: FEC0/10
A unicast GID used within a collection of subnets
that is unique within that collection (e.g., a data
center or campus) but is not necessarily globally
unique. IB routers must not forward any packets with
either a site-local Source GID or a site-local
Destination GID outside of the site.
c. global:
A unicast GID with a global prefix; an IB router may
use this GID to route packets throughout an
enterprise or internet.
1.2.1.2. Multicast GIDs
The multicast GIDs also parallel the IPv6 multicast addresses. The
IB specification defines the multicast GIDs as follows:
FFxy:<112 bits>
Flag bits:
The nibble, denoted by x above, are the 4 flag bits: 000T.
The first 3 bits are reserved and are set to zero. The last
bit is defined as follows:
T=0: denotes a permanently assigned, that is, well-known GID
T=1: denotes a transient group
Scope bits:
The 4 bits, denoted by y in the GID above, are the scope bits.
These scope values are described in Table 1.
scope value Address value
0 Reserved
1 Unassigned
2 Link-local
3 Unassigned
4 Unassigned
5 Site-local
6 Unassigned
7 Unassigned
8 Organization-local
9 Unassigned
0xA Unassigned
0xB Unassigned
0xC Unassigned
0xD Unassigned
0xE Global
0xF Reserved
Table 1
The IB specification further refers to RFC 2373 and RFC 2375 while
defining the well-known multicast addresses. However, it then states
that the well-known addresses apply to IB raw IPv6 datagrams only.
It must be noted though that a multicast group can be associated with
only a single Multicast Global Identifier (MGID). Thus the same MGID
cannot be associated with the UD mode and the Raw Datagram mode.
1.3. InfiniBand Multicast Group Management
IB multicast groups, identified by MGIDs, are managed by the SM. The
SM explicitly programs the IB switches in the fabric to ensure that
the packets are received by all the members of the multicast group
that request the reception of packets. The SM also needs to program
the switches such that packets transmitted to the group by any group
member reach all receivers in the multicast group.
IBA distinguishes between multicast senders and receivers. Though
all members of a multicast group can transmit to the group (and
expect their packets to be correctly forwarded), not all members of
the group are receivers. A port needs to explicitly request that
multicast packets addressed to the group be forwarded to it.
A multicast group is created by sending a join request to the SM. As
will be explained later, IBA defines multiple modes for joining a
multicast group. The subnet manager records the group’s multicast
GID and the associated characteristics. The group characteristics
are defined by the group path MTU, whether the group will be used for
raw datagrams or unreliable datagrams, the service level, the
partition key associated with the group, the Local Identifier (LID)
associated with the group, and so on. These characteristics are
defined at the time of the group creation. The interested reader may
look up the ’MCMemberRecord’ attribute in the IB architecture
specification [IB_ARCH] for the complete list of characteristics that
define a group.
A LID is associated with the multicast group by the SM at the time of
the multicast group creation. The SM determines the multicast tree
based on all the group members and programs the relevant switches.
The Multicast LID (MLID) is used by the switches to route the
packets.
Any member IB port wanting to participate in the multicast group must
join the group. As part of the join operation, the node receives the
group characteristics from the SM. At the same time, the subnet
manager ensures that the requester can indeed participate in the
group by verifying that it can support the group MTU and its
accessibility to the rest of the group members. Other group
characteristics may need verification too.
The SM, for groups that span IB subnet boundaries, must interact with
IB routers to determine the presence of this group in other IB
subnets. If present, the MTU must match across the IB subnets.
P_Key is another characteristic that must match across IB subnets
since the P_Key inserted into a packet is not modified by the IB
switches or IB routers. Thus, if the P_Keys did not match the IB
router(s) itself might drop the packets or destinations on other
subnets might drop the packets.
A join operation may cause the SM to reprogram the fabric so that the
new member can participate in the multicast group. By the same
token, a leave may cause the SM to reprogram the fabric to stop
forwarding the packets to the requester.
1.3.1. Multicast Member Record
The multicast group is maintained by the SM with each of the group
members represented by an MCMemberRecord [IB_ARCH]. Some of its
components are the following:
MGID - Multicast GID for this multicast group
PortGID - Valid GID of the port joining this multicast group
Q_Key - Q_Key to be used by this multicast group
MLID - Multicast LID for this multicast group
MTU - MTU for this multicast group
P_Key - Partition key for this multicast group
SL - Service level for this multicast group
Scope - Same as MGID address scope
JoinState - Join/Leave status requested by the port:
bit 0: FullMember
bit 1: NonMember
bit 2: SendOnlyNonMember
1.3.1.1. JoinState
The JoinState indicates the membership qualities a port wishes to add