recovery of the remaining (S - s) LSP, are based on local policy.
4.16. Recovery Schemes Related Time and Durations
This section gives several typical timing definitions that are of
importance for recovery schemes.
A. Detection time:
The time between the occurrence of the fault or degradation and its
detection. Note that this is a rather theoretical time because, in
practice, this is difficult to measure.
B. Correlation time:
The time between the detection of the fault or degradation and the
reporting of the signal fail or degrade. This time is typically used
in correlating related failures or degradations.
C. Notification time:
The time between the reporting of the signal fail or degrade and the
reception of the indication of this event by the entities that decide
on the recovery switching operation(s).
D. Recovery Switching time:
The time between the initialization of the recovery switching
operation and the moment the normal traffic is selected from the
recovery LSP/span.
E. Total Recovery time:
The total recovery time is defined as the sum of the detection, the
correlation, the notification, and the recovery switching time.
F. Wait To Restore time:
A period of time that must elapse after a recovered fault before an
LSP/span can be used again to transport the normal traffic and/or to
select the normal traffic from.
Note: the hold-off time is defined as the time between the reporting
of signal fail or degrade, and the initialization of the recovery
switching operation. This is useful when multiple layers of recovery
are being used.
4.17. Impairment
A defect or performance degradation, which may lead to SF or SD
trigger.
4.18. Recovery Ratio
The quotient of the actual recovery bandwidth divided by the traffic
bandwidth that is intended to be protected.
4.19. Hitless Protection Switch-over
Protection switch-over, which does not cause data loss, data
duplication, data disorder, or bit errors upon recovery switching
action.
4.20. Network Survivability
The set of capabilities that allows a network to restore affected
traffic in the event of a failure. The degree of survivability is
determined by the network’s capability to survive single and multiple
failures.
4.21. Survivable Network
A network that is capable of restoring traffic in the event of a
failure.
4.22. Escalation
A network survivability action caused by the impossibility of the
survivability function in lower layers.
5. Recovery Phases
It is commonly accepted that recovery implies that the following
generic operations need to be performed when an LSP/span or a node
failure occurs:
- Phase 1: Failure Detection
The action of detecting the impairment (defect of performance
degradation) as a defect condition and the consequential activation
of SF or SD trigger to the control plane (through internal interface
with the transport plane). Thus, failure detection (which should
occur at the transport layer closest to the failure) is the only
phase that cannot be achieved by the control plane alone.
- Phase 2: Failure Localization (and Isolation)
Failure localization provides, to the deciding entity, information
about the location (and thus the identity) of the transport plane
entity that causes the LSP(s)/span(s) failure. The deciding entity
can then make an accurate decision to achieve finer grained recovery
switching action(s).
- Phase 3: Failure Notification
Failure notification phase is used 1) to inform intermediate nodes
that LSP(s)/span(s) failure has occurred and has been detected and 2)
to inform the recovery deciding entities (which can correspond to any
intermediate or end-point of the failed LSP/span) that the
corresponding LSP/span is not available.
- Phase 4: Recovery (Protection or Restoration)
See Section 4.3.
- Phase 5: Reversion (Normalization)
See Section 4.11.
The combination of Failure Detection and Failure Localization and
Notification is referred to as Fault Management.
5.1. Entities Involved During Recovery
The entities involved during the recovery operations can be defined
as follows; these entities are parts of ingress, egress, and
intermediate nodes, as defined previously:
A. Detecting Entity (Failure Detection):
An entity that detects a failure or group of failures; thus providing
a non-correlated list of failures.
B. Reporting Entity (Failure Correlation and Notification):
An entity that can make an intelligent decision on fault correlation
and report the failure to the deciding entity. Fault reporting can
be automatically performed by the deciding entity detecting the
failure.
C. Deciding Entity (part of the failure recovery decision process):
An entity that makes the recovery decision or selects the recovery
resources. This entity communicates the decision to the impacted
LSPs/spans with the recovery actions to be performed.
D. Recovering Entity (part of the failure recovery activation
process):
An entity that participates in the recovery of the LSPs/spans.
The process of moving failed LSPs from a failed (working) span to a
protection span must be initiated by one of the nodes that terminates
the span, e.g., A or B. The deciding (and recovering) entity is
referred to as the "master", while the other node is called the
"slave" and corresponds to a recovering only entity.
Note: The determination of the master and the slave may be based on
configured information or protocol-specific requirements.
6. Protection Schemes
This section clarifies the multiple possible protection schemes and
the specific terminology for the protection.
6.1. 1+1 Protection
1+1 protection has one working LSP/span, one protection LSP/span, and
a permanent bridge. At the ingress node, the normal traffic is
permanently bridged to both the working and protection LSP/span. At
the egress node, the normal traffic is selected from the better of
the two LSPs/spans.
Due to the permanent bridging, the 1+1 protection does not allow an
unprotected extra traffic signal to be provided.
6.2. 1:N (N >= 1) Protection
1:N protection has N working LSPs/spans that carry normal traffic and
1 protection LSP/span that may carry extra-traffic.
At the ingress, the normal traffic is either permanently connected to
its working LSP/span and may be connected to the protection LSP/span
(case of broadcast bridge), or is connected to either its working
LSP/span or the protection LSP/span (case of selector bridge). At
the egress node, the normal traffic is selected from either its
working or protection LSP/span.
Unprotected extra traffic can be transported over the protection
LSP/span whenever the protection LSP/span is not used to carry a
normal traffic.
6.3. M:N (M, N > 1, N >= M) Protection
M:N protection has N working LSPs/spans carrying normal traffic and M
protection LSP/span that may carry extra-traffic.
At the ingress, the normal traffic is either permanently connected to
its working LSP/span and may be connected to one of the protection
LSPs/spans (case of broadcast bridge), or is connected to either its
working LSP/span or one of the protection LSPs/spans (case of
selector bridge). At the egress node, the normal traffic is selected
from either its working or one of the protection LSP/span.
Unprotected extra traffic can be transported over the M protection
LSP/span whenever the protection LSPs/spans is not used to carry a
normal traffic.
6.4. Notes on Protection Schemes
All protection types are either uni- or bi-directional; obviously,
the latter applies only to bi-directional LSPs/spans and requires
coordination between the ingress and egress node during protection
switching.
All protection types except 1+1 unidirectional protection switching
require a communication channel between the ingress and the egress
node.
In the GMPLS context, span protection refers to the full or partial
span recovery of the LSPs carried over that span (see Section 4.15).
7. Restoration Schemes
This section clarifies the multiple possible restoration schemes and
the specific terminology for the restoration.
7.1. Pre-Planned LSP Restoration
Also referred to as pre-planned LSP re-routing. Before failure
detection and/or notification, one or more restoration LSPs are
instantiated between the same ingress-egress node pair as the working
LSP. Note that the restoration resources must be pre-computed, must
be signaled, and may be selected a priori, but may not cross-
connected. Thus, the restoration LSP is not able to carry any
extra-traffic.
The complete establishment of the restoration LSP (i.e., activation)
occurs only after failure detection and/or notification of the
working LSP and requires some additional restoration signaling.
Therefore, this mechanism protects against working LSP failure(s) but
requires activation of the restoration LSP after failure occurrence.
After the ingress node has activated the restoration LSP, the latter
can carry the normal traffic.
Note: when each working LSP is recoverable by exactly one restoration
LSP, one refers also to 1:1 (pre-planned) re-routing without extra-
traffic.
7.1.1. Shared-Mesh Restoration
"Shared-mesh" restoration is defined as a particular case of pre-
planned LSP re-routing that reduces the restoration resource
requirements by allowing multiple restoration LSPs (initiated from
distinct ingress nodes) to share common resources (including links
and nodes.)
7.2. LSP Restoration
Also referred to as LSP re-routing. The ingress node switches the
normal traffic to an alternate LSP that is signaled and fully
established (i.e., cross-connected) after failure detection and/or
notification. The alternate LSP path may be computed after failure
detection and/or notification. In this case, one also refers to
"Full LSP Re-routing."
The alternate LSP is signaled from the ingress node and may reuse the
intermediate node’s resources of the working LSP under failure
condition (and may also include additional intermediate nodes.)
7.2.1. Hard LSP Restoration
Also referred to as hard LSP re-routing. A re-routing operation
where the LSP is released before the full establishment of an
alternate LSP (i.e., break-before-make).
7.2.2. Soft LSP Restoration
Also referred to as soft LSP re-routing. A re-routing operation
where the LSP is released after the full establishment of an
alternate LSP (i.e., make-before-break).
8. Security Considerations
Security considerations are detailed in [RFC4428] and [RFC4426].
9. References
9.1. Normative References
[RFC2119] Bradner, S., "Key words for use in RFCs to Indicate
Requirement Levels", BCP 14, RFC 2119, March 1997.
9.2. Informative References
[RFC3386] Lai, W. and D. McDysan, "Network Hierarchy and
Multilayer Survivability", RFC 3386, November 2002.
[RFC3945] Mannie, E., "Generalized Multi-Protocol Label Switching
(GMPLS) Architecture", RFC 3945, October 2004.
[RFC4426] Lang, J., Rajagopalan B., and D.Papadimitriou, Editors,
"Generalized Multiprotocol Label Switching (GMPLS)
Recovery Functional Specification", RFC 4426, March
2006.
[RFC4428] Papadimitriou D. and E.Mannie, Editors, "Analysis of
Generalized Multi-Protocol Label Switching (GMPLS)-based
Recovery Mechanisms (including Protection and
Restoration)", RFC 4428, March 2006.
For information on the availability of the following documents,
please see http://www.itu.int
[G.808.1] ITU-T, "Generic Protection Switching - Linear trail and
subnetwork protection," Recommendation G.808.1, December
2003.
[G.841] ITU-T, "Types and Characteristics of SDH Network
Protection Architectures," Recommendation G.841, October
1998.
10. Acknowledgements
Many thanks to Adrian Farrel for having thoroughly review this
document.
Editors’ Addresses
Eric Mannie
Perceval
Rue Tenbosch, 9
1000 Brussels
Belgium
Phone: +32-2-6409194
EMail: eric.mannie@perceval.net
Dimitri Papadimitriou
Alcatel
Francis Wellesplein, 1
B-2018 Antwerpen, Belgium
Phone: +32 3 240-8491
EMail: dimitri.papadimitriou@alcatel.be
Full Copyright Statement
Copyright (C) The Internet Society (2006).
This document is subject to the rights, licenses and restrictions
contained in BCP 78, and except as set forth therein, the authors
retain all their rights.
This document and the information contained herein are provided on an
"AS IS" basis and THE CONTRIBUTOR, THE ORGANIZATION HE/SHE REPRESENTS
OR IS SPONSORED BY (IF ANY), THE INTERNET SOCIETY AND THE INTERNET
ENGINEERING TASK FORCE DISCLAIM ALL WARRANTIES, EXPRESS OR IMPLIED,
INCLUDING BUT NOT LIMITED TO ANY WARRANTY THAT THE USE OF THE
INFORMATION HEREIN WILL NOT INFRINGE ANY RIGHTS OR ANY IMPLIED
WARRANTIES OF MERCHANTABILITY OR FITNESS FOR A PARTICULAR PURPOSE.
Intellectual Property
The IETF takes no position regarding the validity or scope of any
Intellectual Property Rights or other rights that might be claimed to
pertain to the implementation or use of the technology described in
this document or the extent to which any license under such rights
might or might not be available; nor does it represent that it has
made any independent effort to identify any such rights. Information
on the procedures with respect to rights in RFC documents can be
found in BCP 78 and BCP 79.
Copies of IPR disclosures made to the IETF Secretariat and any
assurances of licenses to be made available, or the result of an
attempt made to obtain a general license or permission for the use of
such proprietary rights by implementers or users of this
specification can be obtained from the IETF on-line IPR repository at
http://www.ietf.org/ipr.
The IETF invites any interested party to bring to its attention any
copyrights, patents or patent applications, or other proprietary
rights that may cover technology that may be required to implement
this standard. Please address the information to the IETF at
ietf-ipr@ietf.org.
Acknowledgement
Funding for the RFC Editor function is provided by the IETF
Administrative Support Activity (IASA).