Request for Comments: 4251 SSH Communications Security Corp
Category: Standards Track C. Lonvick, Ed.
Cisco Systems, Inc.
January 2006
The Secure Shell (SSH) Protocol Architecture
Status of This Memo
This document specifies an Internet standards track protocol for the
Internet community, and requests discussion and suggestions for
improvements. Please refer to the current edition of the "Internet
Official Protocol Standards" (STD 1) for the standardization state
and status of this protocol. Distribution of this memo is unlimited.
Copyright Notice
Copyright (C) The Internet Society (2006).
Abstract
The Secure Shell (SSH) Protocol is a protocol for secure remote login
and other secure network services over an insecure network. This
document describes the architecture of the SSH protocol, as well as
the notation and terminology used in SSH protocol documents. It also
discusses the SSH algorithm naming system that allows local
extensions. The SSH protocol consists of three major components: The
Transport Layer Protocol provides server authentication,
confidentiality, and integrity with perfect forward secrecy. The
User Authentication Protocol authenticates the client to the server.
The Connection Protocol multiplexes the encrypted tunnel into several
logical channels. Details of these protocols are described in
separate documents.
Table of Contents
1. Introduction ....................................................3
2. Contributors ....................................................3
3. Conventions Used in This Document ...............................4
4. Architecture ....................................................4
4.1. Host Keys ..................................................4
4.2. Extensibility ..............................................6
4.3. Policy Issues ..............................................6
4.4. Security Properties ........................................7
4.5. Localization and Character Set Support .....................7
5. Data Type Representations Used in the SSH Protocols .............8
6. Algorithm and Method Naming ....................................10
7. Message Numbers ................................................11
8. IANA Considerations ............................................12
9. Security Considerations ........................................13
9.1. Pseudo-Random Number Generation ...........................13
9.2. Control Character Filtering ...............................14
9.3. Transport .................................................14
9.3.1. Confidentiality ....................................14
9.3.2. Data Integrity .....................................16
9.3.3. Replay .............................................16
9.3.4. Man-in-the-middle ..................................17
9.3.5. Denial of Service ..................................19
9.3.6. Covert Channels ....................................20
9.3.7. Forward Secrecy ....................................20
9.3.8. Ordering of Key Exchange Methods ...................20
9.3.9. Traffic Analysis ...................................21
9.4. Authentication Protocol ...................................21
9.4.1. Weak Transport .....................................21
9.4.2. Debug Messages .....................................22
9.4.3. Local Security Policy ..............................22
9.4.4. Public Key Authentication ..........................23
9.4.5. Password Authentication ............................23
9.4.6. Host-Based Authentication ..........................23
9.5. Connection Protocol .......................................24
9.5.1. End Point Security .................................24
9.5.2. Proxy Forwarding ...................................24
9.5.3. X11 Forwarding .....................................24
10. References ....................................................26
10.1. Normative References .....................................26
10.2. Informative References ...................................26
Authors’ Addresses ................................................29
Trademark Notice ..................................................29
1. Introduction
Secure Shell (SSH) is a protocol for secure remote login and other
secure network services over an insecure network. It consists of
three major components:
o The Transport Layer Protocol [SSH-TRANS] provides server
authentication, confidentiality, and integrity. It may optionally
also provide compression. The transport layer will typically be
run over a TCP/IP connection, but might also be used on top of any
other reliable data stream.
o The User Authentication Protocol [SSH-USERAUTH] authenticates the
client-side user to the server. It runs over the transport layer
protocol.
o The Connection Protocol [SSH-CONNECT] multiplexes the encrypted
tunnel into several logical channels. It runs over the user
authentication protocol.
The client sends a service request once a secure transport layer
connection has been established. A second service request is sent
after user authentication is complete. This allows new protocols to
be defined and coexist with the protocols listed above.
The connection protocol provides channels that can be used for a wide
range of purposes. Standard methods are provided for setting up
secure interactive shell sessions and for forwarding ("tunneling")
arbitrary TCP/IP ports and X11 connections.
2. Contributors
The major original contributors of this set of documents have been:
Tatu Ylonen, Tero Kivinen, Timo J. Rinne, Sami Lehtinen (all of SSH
Communications Security Corp), and Markku-Juhani O. Saarinen
(University of Jyvaskyla). Darren Moffat was the original editor of
this set of documents and also made very substantial contributions.
Many people contributed to the development of this document over the
years. People who should be acknowledged include Mats Andersson, Ben
Harris, Bill Sommerfeld, Brent McClure, Niels Moller, Damien Miller,
Derek Fawcus, Frank Cusack, Heikki Nousiainen, Jakob Schlyter, Jeff
Van Dyke, Jeffrey Altman, Jeffrey Hutzelman, Jon Bright, Joseph
Galbraith, Ken Hornstein, Markus Friedl, Martin Forssen, Nicolas
Williams, Niels Provos, Perry Metzger, Peter Gutmann, Simon
Josefsson, Simon Tatham, Wei Dai, Denis Bider, der Mouse, and
Tadayoshi Kohno. Listing their names here does not mean that they
endorse this document, but that they have contributed to it.
3. Conventions Used in This Document
All documents related to the SSH protocols shall use the keywords
"MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD",
"SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" to describe
requirements. These keywords are to be interpreted as described in
[RFC2119].
The keywords "PRIVATE USE", "HIERARCHICAL ALLOCATION", "FIRST COME
FIRST SERVED", "EXPERT REVIEW", "SPECIFICATION REQUIRED", "IESG
APPROVAL", "IETF CONSENSUS", and "STANDARDS ACTION" that appear in
this document when used to describe namespace allocation are to be
interpreted as described in [RFC2434].
Protocol fields and possible values to fill them are defined in this
set of documents. Protocol fields will be defined in the message
definitions. As an example, SSH_MSG_CHANNEL_DATA is defined as
follows.
byte SSH_MSG_CHANNEL_DATA
uint32 recipient channel
string data
Throughout these documents, when the fields are referenced, they will
appear within single quotes. When values to fill those fields are
referenced, they will appear within double quotes. Using the above
example, possible values for ’data’ are "foo" and "bar".
4. Architecture
4.1. Host Keys
Each server host SHOULD have a host key. Hosts MAY have multiple
host keys using multiple different algorithms. Multiple hosts MAY
share the same host key. If a host has keys at all, it MUST have at
least one key that uses each REQUIRED public key algorithm (DSS
[FIPS-186-2]).
The server host key is used during key exchange to verify that the
client is really talking to the correct server. For this to be
possible, the client must have a priori knowledge of the server’s
public host key.
Two different trust models can be used:
o The client has a local database that associates each host name (as
typed by the user) with the corresponding public host key. This
method requires no centrally administered infrastructure, and no
third-party coordination. The downside is that the database of
name-to-key associations may become burdensome to maintain.
o The host name-to-key association is certified by a trusted
certification authority (CA). The client only knows the CA root
key, and can verify the validity of all host keys certified by
accepted CAs.
The second alternative eases the maintenance problem, since ideally
only a single CA key needs to be securely stored on the client. On
the other hand, each host key must be appropriately certified by a
central authority before authorization is possible. Also, a lot of
trust is placed on the central infrastructure.
The protocol provides the option that the server name - host key
association is not checked when connecting to the host for the first
time. This allows communication without prior communication of host
keys or certification. The connection still provides protection
against passive listening; however, it becomes vulnerable to active
man-in-the-middle attacks. Implementations SHOULD NOT normally allow
such connections by default, as they pose a potential security
problem. However, as there is no widely deployed key infrastructure
available on the Internet at the time of this writing, this option
makes the protocol much more usable during the transition time until
such an infrastructure emerges, while still providing a much higher
level of security than that offered by older solutions (e.g., telnet
[RFC0854] and rlogin [RFC1282]).
Implementations SHOULD try to make the best effort to check host
keys. An example of a possible strategy is to only accept a host key
without checking the first time a host is connected, save the key in
a local database, and compare against that key on all future
connections to that host.
Implementations MAY provide additional methods for verifying the
correctness of host keys, e.g., a hexadecimal fingerprint derived
from the SHA-1 hash [FIPS-180-2] of the public key. Such
fingerprints can easily be verified by using telephone or other
external communication channels.
All implementations SHOULD provide an option not to accept host keys
that cannot be verified.
The members of this Working Group believe that ’ease of use’ is
critical to end-user acceptance of security solutions, and no
improvement in security is gained if the new solutions are not used.
Thus, providing the option not to check the server host key is
believed to improve the overall security of the Internet, even though
it reduces the security of the protocol in configurations where it is
allowed.
4.2. Extensibility
We believe that the protocol will evolve over time, and some
organizations will want to use their own encryption, authentication,
and/or key exchange methods. Central registration of all extensions
is cumbersome, especially for experimental or classified features.
On the other hand, having no central registration leads to conflicts
in method identifiers, making interoperability difficult.
We have chosen to identify algorithms, methods, formats, and
extension protocols with textual names that are of a specific format.
DNS names are used to create local namespaces where experimental or
classified extensions can be defined without fear of conflicts with
other implementations.
One design goal has been to keep the base protocol as simple as
possible, and to require as few algorithms as possible. However, all
implementations MUST support a minimal set of algorithms to ensure
interoperability (this does not imply that the local policy on all
hosts would necessarily allow these algorithms). The mandatory
algorithms are specified in the relevant protocol documents.
Additional algorithms, methods, formats, and extension protocols can
be defined in separate documents. See Section 6, Algorithm Naming,
for more information.
4.3. Policy Issues
The protocol allows full negotiation of encryption, integrity, key
exchange, compression, and public key algorithms and formats.
Encryption, integrity, public key, and compression algorithms can be
different for each direction.
The following policy issues SHOULD be addressed in the configuration
mechanisms of each implementation:
o Encryption, integrity, and compression algorithms, separately for
each direction. The policy MUST specify which is the preferred
algorithm (e.g., the first algorithm listed in each category).
o Public key algorithms and key exchange method to be used for host
authentication. The existence of trusted host keys for different
public key algorithms also affects this choice.
o The authentication methods that are to be required by the server
for each user. The server’s policy MAY require multiple
authentication for some or all users. The required algorithms MAY
depend on the location from where the user is trying to gain
access.
o The operations that the user is allowed to perform using the
connection protocol. Some issues are related to security; for
example, the policy SHOULD NOT allow the server to start sessions
or run commands on the client machine, and MUST NOT allow
connections to the authentication agent unless forwarding such
connections has been requested. Other issues, such as which
TCP/IP ports can be forwarded and by whom, are clearly issues of
local policy. Many of these issues may involve traversing or
bypassing firewalls, and are interrelated with the local security
policy.
4.4. Security Properties
The primary goal of the SSH protocol is to improve security on the
Internet. It attempts to do this in a way that is easy to deploy,
even at the cost of absolute security.
o All encryption, integrity, and public key algorithms used are
well-known, well-established algorithms.
o All algorithms are used with cryptographically sound key sizes
that are believed to provide protection against even the strongest
cryptanalytic attacks for decades.
o All algorithms are negotiated, and in case some algorithm is
broken, it is easy to switch to some other algorithm without
modifying the base protocol.
Specific concessions were made to make widespread, fast deployment
easier. The particular case where this comes up is verifying that
the server host key really belongs to the desired host; the protocol
allows the verification to be left out, but this is NOT RECOMMENDED.
This is believed to significantly improve usability in the short
term, until widespread Internet public key infrastructures emerge.
4.5. Localization and Character Set Support
For the most part, the SSH protocols do not directly pass text that
would be displayed to the user. However, there are some places where
such data might be passed. When applicable, the character set for
the data MUST be explicitly specified. In most places, ISO-10646
UTF-8 encoding is used [RFC3629]. When applicable, a field is also
provided for a language tag [RFC3066].
One big issue is the character set of the interactive session. There
is no clear solution, as different applications may display data in
different formats. Different types of terminal emulation may also be
employed in the client, and the character set to be used is
effectively determined by the terminal emulation. Thus, no place is
provided for directly specifying the character set or encoding for
terminal session data. However, the terminal emulation type (e.g.,
"vt100") is transmitted to the remote site, and it implicitly
specifies the character set and encoding. Applications typically use
the terminal type to determine what character set they use, or the
character set is determined using some external means. The terminal
emulation may also allow configuring the default character set. In
any case, the character set for the terminal session is considered
primarily a client local issue.
Internal names used to identify algorithms or protocols are normally
never displayed to users, and must be in US-ASCII.
The client and server user names are inherently constrained by what
the server is prepared to accept. They might, however, occasionally
be displayed in logs, reports, etc. They MUST be encoded using ISO
10646 UTF-8, but other encodings may be required in some cases. It
is up to the server to decide how to map user names to accepted user
names. Straight bit-wise, binary comparison is RECOMMENDED.
For localization purposes, the protocol attempts to minimize the
number of textual messages transmitted. When present, such messages
typically relate to errors, debugging information, or some externally
configured data. For data that is normally displayed, it SHOULD be
possible to fetch a localized message instead of the transmitted
message by using a numerical code. The remaining messages SHOULD be
configurable.
5. Data Type Representations Used in the SSH Protocols
byte
A byte represents an arbitrary 8-bit value (octet). Fixed length
data is sometimes represented as an array of bytes, written
byte[n], where n is the number of bytes in the array.
boolean
A boolean value is stored as a single byte. The value 0
represents FALSE, and the value 1 represents TRUE. All non-zero
values MUST be interpreted as TRUE; however, applications MUST NOT
store values other than 0 and 1.
uint32
Represents a 32-bit unsigned integer. Stored as four bytes in the
order of decreasing significance (network byte order). For
example: the value 699921578 (0x29b7f4aa) is stored as 29 b7 f4
aa.
uint64
Represents a 64-bit unsigned integer. Stored as eight bytes in
the order of decreasing significance (network byte order).
string
Arbitrary length binary string. Strings are allowed to contain
arbitrary binary data, including null characters and 8-bit
characters. They are stored as a uint32 containing its length
(number of bytes that follow) and zero (= empty string) or more
bytes that are the value of the string. Terminating null
characters are not used.
Strings are also used to store text. In that case, US-ASCII is
used for internal names, and ISO-10646 UTF-8 for text that might
be displayed to the user. The terminating null character SHOULD
NOT normally be stored in the string. For example: the US-ASCII
string "testing" is represented as 00 00 00 07 t e s t i n g. The
UTF-8 mapping does not alter the encoding of US-ASCII characters.
mpint
Represents multiple precision integers in two’s complement format,
stored as a string, 8 bits per byte, MSB first. Negative numbers
have the value 1 as the most significant bit of the first byte of
the data partition. If the most significant bit would be set for
a positive number, the number MUST be preceded by a zero byte.
Unnecessary leading bytes with the value 0 or 255 MUST NOT be
included. The value zero MUST be stored as a string with zero
bytes of data.
By convention, a number that is used in modular computations in
Z_n SHOULD be represented in the range 0 <= x < n.
Examples:
value (hex) representation (hex)
----------- --------------------
0 00 00 00 00
9a378f9b2e332a7 00 00 00 08 09 a3 78 f9 b2 e3 32 a7
80 00 00 00 02 00 80
-1234 00 00 00 02 ed cc
-deadbeef 00 00 00 05 ff 21 52 41 11
name-list
A string containing a comma-separated list of names. A name-list
is represented as a uint32 containing its length (number of bytes
that follow) followed by a comma-separated list of zero or more
names. A name MUST have a non-zero length, and it MUST NOT
contain a comma (","). As this is a list of names, all of the
elements contained are names and MUST be in US-ASCII. Context may
impose additional restrictions on the names. For example, the
names in a name-list may have to be a list of valid algorithm
identifiers (see Section 6 below), or a list of [RFC3066] language
tags. The order of the names in a name-list may or may not be
significant. Again, this depends on the context in which the list
is used. Terminating null characters MUST NOT be used, neither
for the individual names, nor for the list as a whole.
Examples:
value representation (hex)
----- --------------------
(), the empty name-list 00 00 00 00
("zlib") 00 00 00 04 7a 6c 69 62
("zlib,none") 00 00 00 09 7a 6c 69 62 2c 6e 6f 6e 65
6. Algorithm and Method Naming
The SSH protocols refer to particular hash, encryption, integrity,
compression, and key exchange algorithms or methods by name. There
are some standard algorithms and methods that all implementations
MUST support. There are also algorithms and methods that are defined
in the protocol specification, but are OPTIONAL. Furthermore, it is
expected that some organizations will want to use their own
algorithms or methods.
In this protocol, all algorithm and method identifiers MUST be
printable US-ASCII, non-empty strings no longer than 64 characters.
Names MUST be case-sensitive.
There are two formats for algorithm and method names:
o Names that do not contain an at-sign ("@") are reserved to be
assigned by IETF CONSENSUS. Examples include "3des-cbc", "sha-1",
"hmac-sha1", and "zlib" (the doublequotes are not part of the
name). Names of this format are only valid if they are first
registered with the IANA. Registered names MUST NOT contain an
at-sign ("@"), comma (","), whitespace, control characters (ASCII
codes 32 or less), or the ASCII code 127 (DEL). Names are case-
sensitive, and MUST NOT be longer than 64 characters.
o Anyone can define additional algorithms or methods by using names
in the format name@domainname, e.g., "ourcipher-cbc@example.com".
The format of the part preceding the at-sign is not specified;
however, these names MUST be printable US-ASCII strings, and MUST
NOT contain the comma character (","), whitespace, control
characters (ASCII codes 32 or less), or the ASCII code 127 (DEL).