RFC1341 - MIME (Multipurpose Internet Mail Extensions): Mech(2)

时间:2005-02-14 来源: 作者: 点击:
lines of no more than 76 characters each. All line breaks or other characters not found in Table 1 must be ignored by decoding software. In base64 data, characters other than those in Table 1, line b
  
lines of no more than 76 characters each. All line breaks
or other characters not found in Table 1 must be ignored by
decoding software. In base64 data, characters other than
those in Table 1, line breaks, and other white space
probably indicate a transmission error, about which a
warning message or even a message rejection might be
appropriate under some circumstances.

Special processing is performed if fewer than 24 bits are
available at the end of the data being encoded. A full
encoding quantum is always completed at the end of a body.
When fewer than 24 input bits are available in an input
group, zero bits are added (on the right) to form an
integral number of 6-bit groups. Output character positions
which are not required to represent actual input data are
set to the character "=". Since all base64 input is an
integral number of octets, only the following cases can
arise: (1) the final quantum of encoding input is an
integral multiple of 24 bits; here, the final unit of
encoded output will be an integral multiple of 4 characters
with no "=" padding, (2) the final quantum of encoding input
is exactly 8 bits; here, the final unit of encoded output
will be two characters followed by two "=" padding
characters, or (3) the final quantum of encoding input is
exactly 16 bits; here, the final unit of encoded output will
be three characters followed by one "=" padding character.

Care must be taken to use the proper octets for line breaks
if base64 encoding is applied directly to text material that
has not been converted to canonical form. In particular,
text line breaks should be converted into CRLF sequences

Borenstein & Freed [Page 18]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

prior to base64 encoding. The important thing to note is
that this may be done directly by the encoder rather than in
a prior canonicalization step in some implementations.

NOTE: There is no need to worry about quoting apparent
encapsulation boundaries within base64-encoded parts of
multipart entities because no hyphen characters are used in
the base64 encoding.

6 Additional Optional Content- Header Fields

6.1 Optional Content-ID Header Field

In constructing a high-level user agent, it may be desirable
to allow one body to make reference to another.
Accordingly, bodies may be labeled using the "Content-ID"
header field, which is syntactically identical to the
"Message-ID" header field:

Content-ID := msg-id

Like the Message-ID values, Content-ID values must be
generated to be as unique as possible.

6.2 Optional Content-Description Header Field

The ability to associate some descriptive information with a
given body is often desirable. For example, it may be useful
to mark an "image" body as "a picture of the Space Shuttle
Endeavor." Such text may be placed in the Content-
Description header field.

Content-Description := *text

The description is presumed to be given in the US-ASCII
character set, although the mechanism specified in [RFC-
1342] may be used for non-US-ASCII Content-Description
values.

Borenstein & Freed [Page 19]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

7 The Predefined Content-Type Values

This document defines seven initial Content-Type values and
an extension mechanism for private or experimental types.
Further standard types must be defined by new published
specifications. It is expected that most innovation in new
types of mail will take place as subtypes of the seven types
defined here. The most essential characteristics of the
seven content-types are summarized in Appendix G.

7.1 The Text Content-Type

The text Content-Type is intended for sending material which
is principally textual in form. It is the default Content-
Type. A "charset" parameter may be used to indicate the
character set of the body text. The primary subtype of text
is "plain". This indicates plain (unformatted) text. The
default Content-Type for Internet mail is "text/plain;
charset=us-ascii".

Beyond plain text, there are many formats for representing
what might be known as "extended text" -- text with embedded
formatting and presentation information. An interesting
characteristic of many such representations is that they are
to some extent readable even without the software that
interprets them. It is useful, then, to distinguish them,
at the highest level, from such unreadable data as images,
audio, or text represented in an unreadable form. In the
absence of appropriate interpretation software, it is
reasonable to show subtypes of text to the user, while it is
not reasonable to do so with most nontextual data.

Such formatted textual data should be represented using
subtypes of text. Plausible subtypes of text are typically
given by the common name of the representation format, e.g.,
"text/richtext".

7.1.1 The charset parameter

A critical parameter that may be specified in the Content-
Type field for text data is the character set. This is
specified with a "charset" parameter, as in:

Content-type: text/plain; charset=us-ascii

Unlike some other parameter values, the values of the
charset parameter are NOT case sensitive. The default
character set, which must be assumed in the absence of a
charset parameter, is US-ASCII.

An initial list of predefined character set names can be
found at the end of this section. Additional character sets
may be registered with IANA as described in Appendix F,
although the standardization of their use requires the usual

Borenstein & Freed [Page 20]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

IAB review and approval. Note that if the specified
character set includes 8-bit data, a Content-Transfer-
Encoding header field and a corresponding encoding on the
data are required in order to transmit the body via some
mail transfer protocols, such as SMTP.

The default character set, US-ASCII, has been the subject of
some confusion and ambiguity in the past. Not only were
there some ambiguities in the definition, there have been
wide variations in practice. In order to eliminate such
ambiguity and variations in the future, it is strongly
recommended that new user agents explicitly specify a
character set via the Content-Type header field. "US-ASCII"
does not indicate an arbitrary seven-bit character code, but
specifies that the body uses character coding that uses the
exact correspondence of codes to characters specified in
ASCII. National use variations of ISO 646 [ISO-646] are NOT
ASCII and their use in Internet mail is explicitly
discouraged. The omission of the ISO 646 character set is
deliberate in this regard. The character set name of "US-
ASCII" explicitly refers to ANSI X3.4-1986 [US-ASCII] only.
The character set name "ASCII" is reserved and must not be
used for any purpose.

NOTE: RFC821 explicitly specifies "ASCII", and references
an earlier version of the American Standard. Insofar as one
of the purposes of specifying a Content-Type and character
set is to permit the receiver to unambiguously determine how
the sender intended the coded message to be interpreted,
assuming anything other than "strict ASCII" as the default
would risk unintentional and incompatible changes to the
semantics of messages now being transmitted. This also
implies that messages containing characters coded according
to national variations on ISO 646, or using code-switching
procedures (e.g., those of ISO 2022), as well as 8-bit or
multiple octet character encodings MUST use an appropriate
character set specification to be consistent with this
specification.

The complete US-ASCII character set is listed in [US-ASCII].
Note that the control characters including DEL (0-31, 127)
have no defined meaning apart from the combination CRLF
(ASCII values 13 and 10) indicating a new line. Two of the
characters have de facto meanings in wide use: FF (12) often
means "start subsequent text on the beginning of a new
page"; and TAB or HT (9) often (though not always) means
"move the cursor to the next available column after the
current position where the column number is a multiple of 8
(counting the first column as column 0)." Apart from this,
any use of the control characters or DEL in a body must be
part of a private agreement between the sender and
recipient. Such private agreements are discouraged and
should be replaced by the other capabilities of this
document.

Borenstein & Freed [Page 21]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

NOTE: Beyond US-ASCII, an enormous proliferation of
character sets is possible. It is the opinion of the IETF
working group that a large number of character sets is NOT a
good thing. We would prefer to specify a single character
set that can be used universally for representing all of the
world's languages in electronic mail. Unfortunately,
existing practice in several communities seems to point to
the continued use of multiple character sets in the near
future. For this reason, we define names for a small number
of character sets for which a strong constituent base
exists. It is our hope that ISO 10646 or some other
effort will eventually define a single world character set
which can then be specified for use in Internet mail, but in
the advance of that definition we cannot specify the use of
ISO 10646, Unicode, or any other character set whose
definition is, as of this writing, incomplete.

The defined charset values are:

US-ASCII -- as defined in [US-ASCII].

ISO-8859-X -- where "X" is to be replaced, as
necessary, for the parts of ISO-8859 [ISO-
8859]. Note that the ISO 646 character sets
have deliberately been omitted in favor of
their 8859 replacements, which are the
designated character sets for Internet mail.
As of the publication of this document, the
legitimate values for "X" are the digits 1
through 9.

Note that the character set used, if anything other than
US-ASCII, must always be explicitly specified in the
Content-Type field.

No other character set name may be used in Internet mail
without the publication of a formal specification and its
registration with IANA as described in Appendix F, or by
private agreement, in which case the character set name must
begin with "X-".

Implementors are discouraged from defining new character
sets for mail use unless absolutely necessary.

The "charset" parameter has been defined primarily for the
purpose of textual data, and is described in this section
for that reason. However, it is conceivable that non-
textual data might also wish to specify a charset value for
some purpose, in which case the same syntax and values
should be used.

In general, mail-sending software should always use the
"lowest common denominator" character set possible. For
example, if a body contains only US-ASCII characters, it

Borenstein & Freed [Page 22]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

should be marked as being in the US-ASCII character set, not
ISO-8859-1, which, like all the ISO-8859 family of character
sets, is a superset of US-ASCII. More generally, if a
widely-used character set is a subset of another character
set, and a body contains only characters in the widely-used
subset, it should be labeled as being in that subset. This
will increase the chances that the recipient will be able to
view the mail correctly.

7.1.2 The Text/plain subtype

The primary subtype of text is "plain". This indicates
plain (unformatted) text. The default Content-Type for
Internet mail, "text/plain; charset=us-ascii", describes
existing Internet practice, that is, it is the type of body
defined by RFC822.

7.1.3 The Text/richtext subtype

In order to promote the wider interoperability of simple
formatted text, this document defines an extremely simple
subtype of "text", the "richtext" subtype. This subtype was
designed to meet the following criteria:

1. The syntax must be extremely simple to parse,
so that even teletype-oriented mail systems can
easily strip away the formatting information and
leave only the readable text.

2. The syntax must be extensible to allow for new
formatting commands that are deemed essential.

3. The capabilities must be extremely limited, to
ensure that it can represent no more than is
likely to be representable by the user's primary
word processor. While this limits what can be
sent, it increases the likelihood that what is
sent can be properly displayed.

4. The syntax must be compatible with SGML, so
that, with an appropriate DTD (Document Type
Definition, the standard mechanism for defining a
document type using SGML), a general SGML parser
could be made to parse richtext. However, despite
this compatibility, the syntax should be far
simpler than full SGML, so that no SGML knowledge
is required in order to implement it.

The syntax of "richtext" is very simple. It is assumed, at
the top-level, to be in the US-ASCII character set, unless
of course a different charset parameter was specified in the
Content-type field. All characters represent themselves,
with the exception of the "<" character (ASCII 60), which is
used to mark the beginning of a formatting command.

Borenstein & Freed [Page 23]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

Formatting instructions consist of formatting commands
surrounded by angle brackets ("<>", ASCII 60 and 62). Each
formatting command may be no more than 40 characters in
length, all in US-ASCII, restricted to the alphanumeric and
hyphen ("-") characters. Formatting commands may be preceded
by a forward slash or solidus ("/", ASCII 47), making them
negations, and such negations must always exist to balance
the initial opening commands, except as noted below. Thus,
if the formatting command "<bold>" appears at some point,
there must later be a "</bold>" to balance it. There are
only three exceptions to this "balancing" rule: First, the
command "<lt>" is used to represent a literal "<" character.
Second, the command "<nl>" is used to represent a required
line break. (Otherwise, CRLFs in the data are treated as
equivalent to a single SPACE character.) Finally, the
command "<np>" is used to represent a page break. (NOTE:
The 40 character limit on formatting commands does not
include the "<", ">", or "/" characters that might be
attached to such commands.)

Initially defined formatting commands, not all of which will
be implemented by all richtext implementations, include:

Bold -- causes the subsequent text to be in a bold
font.
Italic -- causes the subsequent text to be in an italic
font.
Fixed -- causes the subsequent text to be in a fixed
width font.
Smaller -- causes the subsequent text to be in a
smaller font.
Bigger -- causes the subsequent text to be in a bigger
font.
Underline -- causes the subsequent text to be
underlined.
Center -- causes the subsequent text to be centered.
FlushLeft -- causes the subsequent text to be left
justified.
FlushRight -- causes the subsequent text to be right
justified.
Indent -- causes the subsequent text to be indented at
the left margin.
IndentRight -- causes the subsequent text to be
indented at the right margin.
Outdent -- causes the subsequent text to be outdented
at the left margin.
OutdentRight -- causes the subsequent text to be
outdented at the right margin.
SamePage -- causes the subsequent text to be grouped,
if possible, on one page.
Subscript -- causes the subsequent text to be
interpreted as a subscript.

Borenstein & Freed [Page 24]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

Superscript -- causes the subsequent text to be
interpreted as a superscript.
Heading -- causes the subsequent text to be interpreted
as a page heading.
Footing -- causes the subsequent text to be interpreted
as a page footing.
ISO-8859-X (for any value of X that is legal as a
"charset" parameter) -- causes the subsequent text
to be interpreted as text in the appropriate
character set.
US-ASCII -- causes the subsequent text to be
interpreted as text in the US-ASCII character set.
Excerpt -- causes the subsequent text to be interpreted
as a textual excerpt from another source.
Typically this will be displayed using indentation
and an alternate font, but such decisions are up
to the viewer.
Paragraph -- causes the subsequent text to be
interpreted as a single paragraph, with
appropriate paragraph breaks (typically blank
space) before and after.
Signature -- causes the subsequent text to be
interpreted as a "signature". Some systems may
wish to display signatures in a smaller font or
otherwise set them apart from the main text of the
message.
Comment -- causes the subsequent text to be interpreted
as a comment, and hence not shown to the reader.
No-op -- has no effect on the subsequent text.
lt -- <lt> is replaced by a literal "<" character. No
balancing </lt> is allowed.
nl -- <nl> causes a line break. No balancing </nl> is
allowed.
np -- <np> causes a page break. No balancing </np> is
allowed.

Each positive formatting command affects all subsequent text
until the matching negative formatting command. Such pairs
of formatting commands must be properly balanced and nested.
Thus, a proper way to describe text in bold italics is:

<bold><italic>the-text</italic></bold>

or, alternately,

<italic><bold>the-text</bold></italic>

but, in particular, the following is illegal
richtext:

<bold><italic>the-text</bold></italic>

NOTE: The nesting requirement for formatting commands
imposes a slightly higher burden upon the composers of

Borenstein & Freed [Page 25]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

richtext bodies, but potentially simplifies richtext
displayers by allowing them to be stack-based. The main
goal of richtext is to be simple enough to make multifont,
formatted email widely readable, so that those with the
capability of sending it will be able to do so with
confidence. Thus slightly increased complexity in the
composing software was deemed a reasonable tradeoff for
simplified reading software. Nonetheless, implementors of
richtext readers are encouraged to follow the general
Internet guidelines of being conservative in what you send
and liberal in what you accept. Those implementations that
can do so are encouraged to deal reasonably with improperly
nested richtext.

Implementations must regard any unrecognized formatting
command as equivalent to "No-op", thus facilitating future
extensions to "richtext". Private extensions may be defined
using formatting commands that begin with "X-", by analogy
to Internet mail header field names.

It is worth noting that no special behavior is required for
the TAB (HT) character. It is recommended, however, that, at
least when fixed-width fonts are in use, the common
semantics of the TAB (HT) character should be observed,
namely that it moves to the next column position that is a
multiple of 8. (In other words, if a TAB (HT) occurs in
column n, where the leftmost column is column 0, then that
TAB (HT) should be replaced by 8-(n mod 8) SPACE
characters.)

Richtext also differentiates between "hard" and "soft" line
breaks. A line break (CRLF) in the richtext data stream is
interpreted as a "soft" line break, one that is included
only for purposes of mail transport, and is to be treated as
white space by richtext interpreters. To include a "hard"
line break (one that must be displayed as such), the "<nl>"
or "<paragraph> formatting constructs should be used. In
general, a soft line break should be treated as white space,
but when soft line breaks immediately follow a <nl> or a
</paragraph> tag they should be ignored rather than treated
as white space.

Putting all this together, the following "text/richtext"
body fragment:

<bold>Now</bold> is the time for
<italic>all</italic> good men
<smaller>(and <lt>women>)</smaller> to
<ignoreme></ignoreme> come

to the aid of their
<nl>

Borenstein & Freed [Page 26]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

beloved <nl><nl>country. <comment> Stupid
quote! </comment> -- the end

represents the following formatted text (which will, no
doubt, look cryptic in the text-only version of this
document):

Now is the time for all good men (and <women>) to
come to the aid of their
beloved

country. -- the end

Richtext conformance: A minimal richtext implementation is
one that simply converts "<lt>" to "<", converts CRLFs to
SPACE, converts <nl> to a newline according to local newline
convention, removes everything between a <comment> command
and the next balancing </comment> command, and removes all
other formatting commands (all text enclosed in angle
brackets).

NOTE ON THE RELATIONSHIP OF RICHTEXT TO SGML: Richtext is
decidedly not SGML, and must not be used to transport
arbitrary SGML documents. Those who wish to use SGML
document types as a mail transport format must define a new
text or application subtype, e.g., "text/sgml-dtd-whatever"
or "application/sgml-dtd-whatever", depending on the
perceived readability of the DTD in use. Richtext is
designed to be compatible with SGML, and specifically so
that it will be possible to define a richtext DTD if one is
needed. However, this does not imply that arbitrary SGML
can be called richtext, nor that richtext implementors have
any need to understand SGML; the description in this
document is a complete definition of richtext, which is far
simpler than complete SGML.

NOTE ON THE INTENDED USE OF RICHTEXT: It is recognized that
implementors of future mail systems will want rich text
functionality far beyond that currently defined for
richtext. The intent of richtext is to provide a common
format for expressing that functionality in a form in which
much of it, at least, will be understood by interoperating
software. Thus, in particular, software with a richer
notion of formatted text than richtext can still use
richtext as its basic representation, but can extend it with
new formatting commands and by hiding information specific
to that software system in richtext comments. As such
systems evolve, it is expected that the definition of
richtext will be further refined by future published
specifications, but richtext as defined here provides a
platform on which evolutionary refinements can be based.

IMPLEMENTATION NOTE: In some environments, it might be
impossible to combine certain richtext formatting commands,

Borenstein & Freed [Page 27]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

whereas in others they might be combined easily. For
example, the combination of <bold> and <italic> might
produce bold italics on systems that support such fonts, but
there exist systems that can make text bold or italicized,
but not both. In such cases, the most recently issued
recognized formatting command should be preferred.

One of the major goals in the design of richtext was to make
it so simple that even text-only mailers will implement
richtext-to-plain-text translators, thus increasing the
likelihood that multifont text will become "safe" to use
very widely. To demonstrate this simplicity, an extremely
simple 35-line C program that converts richtext input into
plain text output is included in Appendix D.

Borenstein & Freed [Page 28]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

7.2 The Multipart Content-Type

In the case of multiple part messages, in which one or more
different sets of data are combined in a single body, a
"multipart" Content-Type field must appear in the entity's
header. The body must then contain one or more "body parts,"
each preceded by an encapsulation boundary, and the last one
followed by a closing boundary. Each part starts with an
encapsulation boundary, and then contains a body part
consisting of header area, a blank line, and a body area.
Thus a body part is similar to an RFC822 message in syntax,
but different in meaning.

A body part is NOT to be interpreted as actually being an
RFC822 message. To begin with, NO header fields are
actually required in body parts. A body part that starts
with a blank line, therefore, is allowed and is a body part
for which all default values are to be assumed. In such a
case, the absence of a Content-Type header field implies
that the encapsulation is plain US-ASCII text. The only
header fields that have defined meaning for body parts are
those the names of which begin with "Content-". All other
header fields are generally to be ignored in body parts.
Although they should generally be retained in mail
processing, they may be discarded by gateways if necessary.
Such other fields are permitted to appear in body parts but
should not be depended on. "X-" fields may be created for
experimental or private purposes, with the recognition that
the information they contain may be lost at some gateways.

The distinction between an RFC822 message and a body part
is subtle, but important. A gateway between Internet and
X.400 mail, for example, must be able to tell the difference
between a body part that contains an image and a body part
that contains an encapsulated message, the body of which is
an image. In order to represent the latter, the body part
must have "Content-Type: message", and its body (after the
blank line) must be the encapsulated message, with its own
"Content-Type: image" header field. The use of similar
syntax facilitates the conversion of messages to body parts,
and vice versa, but the distinction between the two must be
understood by implementors. (For the special case in which
all parts actually are messages, a "digest" subtype is also
defined.)

As stated previously, each body part is preceded by an
encapsulation boundary. The encapsulation boundary MUST NOT
appear inside any of the encapsulated parts. Thus, it is
crucial that the composing agent be able to choose and
specify the unique boundary that will separate the parts.

All present and future subtypes of the "multipart" type must
use an identical syntax. Subtypes may differ in their
semantics, and may impose additional restrictions on syntax,

Borenstein & Freed [Page 29]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

but must conform to the required syntax for the multipart
type. This requirement ensures that all conformant user
agents will at least be able to recognize and separate the
parts of any multipart entity, even of an unrecognized
subtype.

As stated in the definition of the Content-Transfer-Encoding
field, no encoding other than "7bit", "8bit", or "binary" is
permitted for entities of type "multipart". The multipart
delimiters and header fields are always 7-bit ASCII in any
case, and data within the body parts can be encoded on a
part-by-part basis, with Content-Transfer-Encoding fields
for each appropriate body part.

Mail gateways, relays, and other mail handling agents are
commonly known to alter the top-level header of an RFC822
message. In particular, they frequently add, remove, or
reorder header fields. Such alterations are explicitly
forbidden for the body part headers embedded in the bodies
of messages of type "multipart."

7.2.1 Multipart: The common syntax

All subtypes of "multipart" share a common syntax, defined
in this section. A simple example of a multipart message
also appears in this section. An example of a more complex
multipart message is given in Appendix C.

The Content-Type field for multipart entities requires one
parameter, "boundary", which is used to specify the
encapsulation boundary. The encapsulation boundary is
defined as a line consisting entirely of two hyphen
characters ("-", decimal code 45) followed by the boundary
parameter value from the Content-Type header field.

NOTE: The hyphens are for rough compatibility with the
earlier RFC934 method of message encapsulation, and for
ease of searching for the boundaries in some
implementations. However, it should be noted that multipart
messages are NOT completely compatible with RFC934
encapsulations; in particular, they do not obey RFC934
quoting conventions for embedded lines that begin with
hyphens. This mechanism was chosen over the RFC934
mechanism because the latter causes lines to grow with each
level of quoting. The combination of this growth with the
fact that SMTP implementations sometimes wrap long lines
made the RFC934 mechanism unsuitable for use in the event
that deeply-nested multipart structuring is ever desired.

Thus, a typical multipart Content-Type header field might
look like this:

Content-Type: multipart/mixed;

Borenstein & Freed [Page 30]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

boundary=gc0p4Jq0M2Yt08jU534c0p

This indicates that the entity consists of several parts,
each itself with a structure that is syntactically identical
to an RFC822 message, except that the header area might be
completely empty, and that the parts are each preceded by
the line

--gc0p4Jq0M2Yt08jU534c0p

Note that the encapsulation boundary must occur at the
beginning of a line, i.e., following a CRLF, and that that
initial CRLF is considered to be part of the encapsulation
boundary rather than part of the preceding part. The
boundary must be followed immediately either by another CRLF
and the header fields for the next part, or by two CRLFs, in
which case there are no header fields for the next part (and
it is therefore assumed to be of Content-Type text/plain).

NOTE: The CRLF preceding the encapsulation line is
considered part of the boundary so that it is possible to
have a part that does not end with a CRLF (line break).
Body parts that must be considered to end with line breaks,
therefore, should have two CRLFs preceding the encapsulation
line, the first of which is part of the preceding body part,
and the second of which is part of the encapsulation
boundary.

The requirement that the encapsulation boundary begins with
a CRLF implies that the body of a multipart entity must
itself begin with a CRLF before the first encapsulation line
-- that is, if the "preamble" area is not used, the entity
headers must be followed by TWO CRLFs. This is indeed how
such entities should be composed. A tolerant mail reading
program, however, may interpret a body of type multipart
that begins with an encapsulation line NOT initiated by a
CRLF as also being an encapsulation boundary, but a
compliant mail sending program must not generate such
entities.

Encapsulation boundaries must not appear within the
encapsulations, and must be no longer than 70 characters,
not counting the two leading hyphens.

The encapsulation boundary following the last body part is a
distinguished delimiter that indicates that no further body
parts will follow. Such a delimiter is identical to the
previous delimiters, with the addition of two more hyphens
at the end of the line:

--gc0p4Jq0M2Yt08jU534c0p--

There appears to be room for additional information prior to
the first encapsulation boundary and following the final

Borenstein & Freed [Page 31]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

boundary. These areas should generally be left blank, and
implementations should ignore anything that appears before
the first boundary or after the last one.

NOTE: These "preamble" and "epilogue" areas are not used
because of the lack of proper typing of these parts and the
lack of clear semantics for handling these areas at
gateways, particularly X.400 gateways.

NOTE: Because encapsulation boundaries must not appear in
the body parts being encapsulated, a user agent must
exercise care to choose a unique boundary. The boundary in
the example above could have been the result of an algorithm
designed to produce boundaries with a very low probability
of already existing in the data to be encapsulated without
having to prescan the data. Alternate algorithms might
result in more 'readable' boundaries for a recipient with an
old user agent, but would require more attention to the
possibility that the boundary might appear in the
encapsulated part. The simplest boundary possible is
something like "---", with a closing boundary of "-----".

As a very simple example, the following multipart message
has two parts, both of them plain text, one of them
explicitly typed and one of them implicitly typed:

From: Nathaniel Borenstein <nsb@bellcore.com>
To: Ned Freed <ned@innosoft.com>
Subject: Sample message
MIME-Version: 1.0
Content-type: multipart/mixed; boundary="simple
boundary"

This is the preamble. It is to be ignored, though it
is a handy place for mail composers to include an
explanatory note to non-MIME compliant readers.
--simple boundary

This is implicitly typed plain ASCII text.
It does NOT end with a linebreak.
--simple boundary
Content-type: text/plain; charset=us-ascii

This is explicitly typed plain ASCII text.
It DOES end with a linebreak.

--simple boundary--
This is the epilogue. It is also to be ignored.

The use of a Content-Type of multipart in a body part within
another multipart entity is explicitly allowed. In such
cases, for obvious reasons, care must be taken to ensure
that each nested multipart entity must use a different
boundary delimiter. See Appendix C for an example of nested

Borenstein & Freed [Page 32]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

multipart entities.

The use of the multipart Content-Type with only a single
body part may be useful in certain contexts, and is
explicitly permitted.

The only mandatory parameter for the multipart Content-Type
is the boundary parameter, which consists of 1 to 70
characters from a set of characters known to be very robust
through email gateways, and NOT ending with white space.
(If a boundary appears to end with white space, the white
space must be presumed to have been added by a gateway, and
should be deleted.) It is formally specified by the
following BNF:

boundary := 0*69<bchars> bcharsnospace

bchars := bcharsnospace / " "

bcharsnospace := DIGIT / ALPHA / "'" / "(" / ")" / "+" /
"_"
/ "," / "-" / "." / "/" / ":" / "=" / "?"

Overall, the body of a multipart entity may be specified as
follows:

multipart-body := preamble 1*encapsulation
close-delimiter epilogue

encapsulation := delimiter CRLF body-part

delimiter := CRLF "--" boundary ; taken from Content-Type
field.
; when content-type is
multipart
; There must be no space
; between "--" and boundary.

close-delimiter := delimiter "--" ; Again, no space before
"--"

preamble := *text ; to be ignored upon
receipt.

epilogue := *text ; to be ignored upon
receipt.

body-part = <"message" as defined in RFC822,
with all header fields optional, and with the
specified delimiter not occurring anywhere in
the message body, either on a line by itself
or as a substring anywhere. Note that the

Borenstein & Freed [Page 33]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

semantics of a part differ from the semantics
of a message, as described in the text.>

NOTE: Conspicuously missing from the multipart type is a
notion of structured, related body parts. In general, it
seems premature to try to standardize interpart structure
yet. It is recommended that those wishing to provide a more
structured or integrated multipart messaging facility should
define a subtype of multipart that is syntactically
identical, but that always expects the inclusion of a
distinguished part that can be used to specify the structure
and integration of the other parts, probably referring to
them by their Content-ID field. If this approach is used,
other implementations will not recognize the new subtype,
but will treat it as the primary subtype (multipart/mixed)
and will thus be able to show the user the parts that are
recognized.

7.2.2 The Multipart/mixed (primary) subtype

The primary subtype for multipart, "mixed", is intended for
use when the body parts are independent and intended to be
displayed serially. Any multipart subtypes that an
implementation does not recognize should be treated as being
of subtype "mixed".

7.2.3 The Multipart/alternative subtype

The multipart/alternative type is syntactically identical to
multipart/mixed, but the semantics are different. In
particular, each of the parts is an "alternative" version of
the same information. User agents should recognize that the
content of the various parts are interchangeable. The user
agent should either choose the "best" type based on the
user's environment and preferences, or offer the user the
available alternatives. In general, choosing the best type
means displaying only the LAST part that can be displayed.
This may be used, for example, to send mail in a fancy text
format in such a way that it can easily be displayed
anywhere:

From: Nathaniel Borenstein <nsb@bellcore.com>
To: Ned Freed <ned@innosoft.com>
Subject: Formatted text mail
MIME-Version: 1.0
Content-Type: multipart/alternative; boundary=boundary42

--boundary42
Content-Type: text/plain; charset=us-ascii

...plain text version of message goes here....

Borenstein & Freed [Page 34]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

--boundary42
Content-Type: text/richtext

.... richtext version of same message goes here ...
--boundary42
Content-Type: text/x-whatever

.... fanciest formatted version of same message goes here
...
--boundary42--

In this example, users whose mail system understood the
"text/x-whatever" format would see only the fancy version,
while other users would see only the richtext or plain text
version, depending on the capabilities of their system.

In general, user agents that compose multipart/alternative
entities should place the body parts in increasing order of
preference, that is, with the preferred format last. For
fancy text, the sending user agent should put the plainest
format first and the richest format last. Receiving user
agents should pick and display the last format they are
capable of displaying. In the case where one of the
alternatives is itself of type "multipart" and contains
unrecognized sub-parts, the user agent may choose either to
show that alternative, an earlier alternative, or both.

NOTE: From an implementor's perspective, it might seem more
sensible to reverse this ordering, and have the plainest
alternative last. However, placing the plainest alternative
first is the friendliest possible option when
mutlipart/alternative entities are viewed using a non-MIME-
compliant mail reader. While this approach does impose some
burden on compliant mail readers, interoperability with
older mail readers was deemed to be more important in this
case.

It may be the case that some user agents, if they can
recognize more than one of the formats, will prefer to offer
the user the choice of which format to view. This makes
sense, for example, if mail includes both a nicely-formatted
image version and an easily-edited text version. What is
most critical, however, is that the user not automatically
be shown multiple versions of the same data. Either the
user should be shown the last recognized version or should
explicitly be given the choice.

Borenstein & Freed [Page 35]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

7.2.4 The Multipart/digest subtype

This document defines a "digest" subtype of the multipart
Content-Type. This type is syntactically identical to
multipart/mixed, but the semantics are different. In
particular, in a digest, the default Content-Type value for
a body part is changed from "text/plain" to
"message/rfc822". This is done to allow a more readable
digest format that is largely compatible (except for the
quoting convention) with RFC934.

A digest in this format might, then, look something like
this:

From: Moderator-Address
MIME-Version: 1.0
Subject: Internet Digest, volume 42
Content-Type: multipart/digest;
boundary="---- next message ----"

------ next message ----

From: someone-else
Subject: my opinion

...body goes here ...

------ next message ----

From: someone-else-again
Subject: my different opinion

... another body goes here...

------ next message ------

7.2.5 The Multipart/parallel subtype

This document defines a "parallel" subtype of the multipart
Content-Type. This type is syntactically identical to
multipart/mixed, but the semantics are different. In
particular, in a parallel entity, all of the parts are
intended to be presented in parallel, i.e., simultaneously,
on hardware and software that are capable of doing so.
Composing agents should be aware that many mail readers will
lack this capability and will show the parts serially in any
event.

Borenstein & Freed [Page 36]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

7.3 The Message Content-Type

It is frequently desirable, in sending mail, to encapsulate
another mail message. For this common operation, a special
Content-Type, "message", is defined. The primary subtype,
message/rfc822, has no required parameters in the Content-
Type field. Additional subtypes, "partial" and "External-
body", do have required parameters. These subtypes are
explained below.

NOTE: It has been suggested that subtypes of message might
be defined for forwarded or rejected messages. However,
forwarded and rejected messages can be handled as multipart
messages in which the first part contains any control or
descriptive information, and a second part, of type
message/rfc822, is the forwarded or rejected message.
Composing rejection and forwarding messages in this manner
will preserve the type information on the original message
and allow it to be correctly presented to the recipient, and
hence is strongly encouraged.

As stated in the definition of the Content-Transfer-Encoding
field, no encoding other than "7bit", "8bit", or "binary" is
permitted for messages or parts of type "message". The
message header fields are always US-ASCII in any case, and
data within the body can still be encoded, in which case the
Content-Transfer-Encoding header field in the encapsulated
message will reflect this. Non-ASCII text in the headers of
an encapsulated message can be specified using the
mechanisms described in [RFC-1342].

Mail gateways, relays, and other mail handling agents are
commonly known to alter the top-level header of an RFC822
message. In particular, they frequently add, remove, or
reorder header fields. Such alterations are explicitly
forbidden for the encapsulated headers embedded in the
bodies of messages of type "message."

7.3.1 The Message/rfc822 (primary) subtype

A Content-Type of "message/rfc822" indicates that the body
contains an encapsulated message, with the syntax of an RFC
822 message.

7.3.2 The Message/Partial subtype

A subtype of message, "partial", is defined in order to
allow large objects to be delivered as several separate
pieces of mail and automatically reassembled by the
receiving user agent. (The concept is similar to IP
fragmentation/reassembly in the basic Internet Protocols.)
This mechanism can be used when intermediate transport
agents limit the size of individual messages that can be
sent. Content-Type "message/partial" thus indicates that

Borenstein & Freed [Page 37]

RFC1341MIME: Multipurpose Internet Mail ExtensionsJune 1992

the body contains a fragment of a larger message.

Three parameters must be specified in the Content-Type field
of type message/partial: The first, "id", is a unique
identifier, as close to a world-unique identifier as
possible, to be used to match the parts together. (In
general, the identifier is essentially a message-id; if
placed in double quotes, it can be any message-id, in
accordance with the BNF for "parameter" given earlier in
this specification.) The second, "number", an integer, is
the part number, which indicates where this part fits into
the sequence of fragments. The third, "total", another
integer, is the total number of parts. This third subfield
is required on the final part, and is optional on the
earlier parts. Note also that these parameters may be given
in any order.

Thus, part 2 of a 3-part message may have either of the
following header fields:

Content-Type: Message/Partial;
number=2; total=3;
id="oc=jpbe0M2Yt4s@thumper.bellcore.com";

Content-Type: Message/Partial;
id="oc=jpbe0M2Yt4s@thumper.bellcore.com";
number=2

But part 3 MUST specify the total number of parts:

Content-Type: Message/Partial;
number=3; total=3;
id="oc=jpbe0M2Yt4s@thumper.bellcore.com";

Note that part numbering begins with 1, not 0.

When the parts of a message broken up in this manner are put
together, the result is a complete RFC822 format message,
which may have its own Content-Type header field, and thus
may contain any other data type.

Message fragmentation and reassembly: The semantics of a
reassembled partial message must be those of the "inner"
message, rather than of a message containing the inner
message. This makes it possible, for example, to send a
large audio message as several partial messages, and still
------分隔线----------------------------
顶一下
(0)
0%
踩一下
(0)
0%
------分隔线----------------------------
最新评论 查看所有评论
发表评论 查看所有评论
请自觉遵守互联网相关的政策法规,严禁发布色情、暴力、反动的言论。
评价:
表情:
用户名: 密码: 验证码:
推荐内容