language subtag, there MUST be an attempt to register the language
with ISO 639. Subtags MUST NOT be registered for codes that exist
in ISO 639-1 or ISO 639-2, that are under consideration by the ISO
639 maintenance or registration authorities, or that have never
been attempted for registration with those authorities. If ISO
639 has previously rejected a language for registration, it is
reasonable to assume that there must be additional, very
compelling evidence of need before it will be registered in the
IANA registry (to the extent that it is very unlikely that any
subtags will be registered of this type).
o Dialect or other divisions or variations within a language, its
orthography, writing system, regional or historical usage,
transliteration or other transformation, or distinguishing
variation MAY be registered as variant subtags. An example is the
’rozaj’ subtag (the Resian dialect of Slovenian).
o The addition or maintenance of fields (generally of an
informational nature) in Tag or Subtag records as described in
Section 3.1 and subject to the stability provisions in
Section 3.4. This includes descriptions, comments, deprecation
and preferred values for obsolete or withdrawn codes, or the
addition of script or extlang information to primary language
subtags.
o The addition of records and related field value changes necessary
to reflect assignments made by ISO 639, ISO 15924, ISO 3166, and
UN M.49 as described in Section 3.4.
Subtags proposed for registration that would cause all or part of a
grandfathered tag to become redundant but whose meaning conflicts
with or alters the meaning of the grandfathered tag MUST be rejected.
This document leaves the decision on what subtags or changes to
subtags are appropriate (or not) to the registration process
described in Section 3.5.
Note: four-character primary language subtags are reserved to allow
for the possibility of alpha4 codes in some future addition to the
ISO 639 family of standards.
ISO 639 defines a maintenance agency for additions to and changes in
the list of languages in ISO 639. This agency is:
International Information Centre for Terminology (Infoterm)
Aichholzgasse 6/12, AT-1120
Wien, Austria
Phone: +43 1 26 75 35 Ext. 312 Fax: +43 1 216 32 72
ISO 639-2 defines a maintenance agency for additions to and changes
in the list of languages in ISO 639-2. This agency is:
Library of Congress
Network Development and MARC Standards Office
Washington, D.C. 20540 USA
Phone: +1 202 707 6237 Fax: +1 202 707 0115
URL: http://www.loc.gov/standards/iso639-2
The maintenance agency for ISO 3166 (country codes) is:
ISO 3166 Maintenance Agency
c/o International Organization for Standardization
Case postale 56
CH-1211 Geneva 20 Switzerland
Phone: +41 22 749 72 33 Fax: +41 22 749 73 49
URL: http://www.iso.org/iso/en/prods-services/iso3166ma/index.html
The registration authority for ISO 15924 (script codes) is:
Unicode Consortium Box 391476
Mountain View, CA 94039-1476, USA
URL: http://www.unicode.org/iso15924
The Statistics Division of the United Nations Secretariat maintains
the Standard Country or Area Codes for Statistical Use and can be
reached at:
Statistical Services Branch
Statistics Division
United Nations, Room DC2-1620
New York, NY 10017, USA
Fax: +1-212-963-0623
E-mail: statistics@un.org
URL: http://unstats.un.org/unsd/methods/m49/m49alpha.htm
3.7. Extensions and Extensions Registry
Extension subtags are those introduced by single-character subtags
("singletons") other than ’x’. They are reserved for the generation
of identifiers that contain a language component and are compatible
with applications that understand language tags.
The structure and form of extensions are defined by this document so
that implementations can be created that are forward compatible with
applications that might be created using singletons in the future.
In addition, defining a mechanism for maintaining singletons will
lend stability to this document by reducing the likely need for
future revisions or updates.
Single-character subtags are assigned by IANA using the "IETF
Consensus" policy defined by [RFC2434]. This policy requires the
development of an RFC, which SHALL define the name, purpose,
processes, and procedures for maintaining the subtags. The
maintaining or registering authority, including name, contact email,
discussion list email, and URL location of the registry, MUST be
indicated clearly in the RFC. The RFC MUST specify or include each
of the following:
o The specification MUST reference the specific version or revision
of this document that governs its creation and MUST reference this
section of this document.
o The specification and all subtags defined by the specification
MUST follow the ABNF and other rules for the formation of tags and
subtags as defined in this document. In particular, it MUST
specify that case is not significant and that subtags MUST NOT
exceed eight characters in length.
o The specification MUST specify a canonical representation.
o The specification of valid subtags MUST be available over the
Internet and at no cost.
o The specification MUST be in the public domain or available via a
royalty-free license acceptable to the IETF and specified in the
RFC.
o The specification MUST be versioned, and each version of the
specification MUST be numbered, dated, and stable.
o The specification MUST be stable. That is, extension subtags,
once defined by a specification, MUST NOT be retracted or change
in meaning in any substantial way.
o The specification MUST include in a separate section the
registration form reproduced in this section (below) to be used in
registering the extension upon publication as an RFC.
o IANA MUST be informed of changes to the contact information and
URL for the specification.
IANA will maintain a registry of allocated single-character
(singleton) subtags. This registry MUST use the record-jar format
described by the ABNF in Section 3.1. Upon publication of an
extension as an RFC, the maintaining authority defined in the RFC
MUST forward this registration form to iesg@ietf.org, who MUST
forward the request to iana@iana.org. The maintaining authority of
the extension MUST maintain the accuracy of the record by sending an
updated full copy of the record to iana@iana.org with the subject
line "LANGUAGE TAG EXTENSION UPDATE" whenever content changes. Only
the ’Comments’, ’Contact_Email’, ’Mailing_List’, and ’URL’ fields MAY
be modified in these updates.
Failure to maintain this record, maintain the corresponding registry,
or meet other conditions imposed by this section of this document MAY
be appealed to the IESG [RFC2028] under the same rules as other IETF
decisions (see [RFC2026]) and MAY result in the authority to maintain
the extension being withdrawn or reassigned by the IESG.
%%
Identifier:
Description:
Comments:
Added:
RFC:
Authority:
Contact_Email:
Mailing_List:
URL:
%%
Figure 6: Format of Records in the Language Tag Extensions Registry
’Identifier’ contains the single-character subtag (singleton)
assigned to the extension. The Internet-Draft submitted to define
the extension SHOULD specify which letter or digit to use, although
the IESG MAY change the assignment when approving the RFC.
’Description’ contains the name and description of the extension.
’Comments’ is an OPTIONAL field and MAY contain a broader description
of the extension.
’Added’ contains the date the RFC was published in the "full-date"
format specified in [RFC3339]. For example: 2004-06-28 represents
June 28, 2004, in the Gregorian calendar.
’RFC’ contains the RFC number assigned to the extension.
’Authority’ contains the name of the maintaining authority for the
extension.
’Contact_Email’ contains the email address used to contact the
maintaining authority.
’Mailing_List’ contains the URL or subscription email address of the
mailing list used by the maintaining authority.
’URL’ contains the URL of the registry for this extension.
The determination of whether an Internet-Draft meets the above
conditions and the decision to grant or withhold such authority rests
solely with the IESG and is subject to the normal review and appeals
process associated with the RFC process.
Extension authors are strongly cautioned that many (including most
well-formed) processors will be unaware of any special relationships
or meaning inherent in the order of extension subtags. Extension
authors SHOULD avoid subtag relationships or canonicalization
mechanisms that interfere with matching or with length restrictions
that sometimes exist in common protocols where the extension is used.
In particular, applications MAY truncate the subtags in doing
matching or in fitting into limited lengths, so it is RECOMMENDED
that the most significant information be in the most significant
(left-most) subtags and that the specification gracefully handle
truncated subtags.
When a language tag is to be used in a specific, known, protocol, it
is RECOMMENDED that the language tag not contain extensions not
supported by that protocol. In addition, note that some protocols
MAY impose upper limits on the length of the strings used to store or
transport the language tag.
3.8. Initialization of the Registries
Upon adoption of this document, an initial version of the Language
Subtag Registry containing the various subtags initially valid in a
language tag is necessary. This collection of subtags, along with a
description of the process used to create it, is described by
[RFC4645]. IANA SHALL publish the initial version of the registry
described by this document from the content of [RFC4645]. Once
published by IANA, the maintenance procedures, rules, and
registration processes described in this document will be available
for new registrations or updates.
Registrations that are in process under the rules defined in
[RFC3066] when this document is adopted MAY be completed under the
former rules, at the discretion of the Language Tag Reviewer (as
described in [RFC3066]). Until the IESG officially appoints a
Language Subtag Reviewer, the existing Language Tag Reviewer SHALL
serve as the Language Subtag Reviewer.
Any new registrations submitted using the RFC 3066 forms or format
after the adoption of this document and publication of the registry
by IANA MUST be rejected.
An initial version of the Language Tag Extensions Registry described
in Section 3.7 is also needed. The Language Tag Extensions Registry
SHALL be initialized with a single record containing a single field
of type "File-Date" as a placeholder for future assignments.
4. Formation and Processing of Language Tags
This section addresses how to use the information in the registry
with the tag syntax to choose, form, and process language tags.
4.1. Choice of Language Tag
One is sometimes faced with the choice between several possible tags
for the same body of text.
Interoperability is best served when all users use the same language
tag in order to represent the same language. If an application has
requirements that make the rules here inapplicable, then that
application risks damaging interoperability. It is strongly
RECOMMENDED that users not define their own rules for language tag
choice.
Subtags SHOULD only be used where they add useful distinguishing
information; extraneous subtags interfere with the meaning,
understanding, and processing of language tags. In particular, users
and implementations SHOULD follow the ’Prefix’ and ’Suppress-Script’
fields in the registry (defined in Section 3.1): these fields provide
guidance on when specific additional subtags SHOULD (and SHOULD NOT)
be used in a language tag.
Of particular note, many applications can benefit from the use of
script subtags in language tags, as long as the use is consistent for
a given context. Script subtags were not formally defined in RFC
3066 and their use can affect matching and subtag identification by
implementations of RFC 3066, as these subtags appear between the
primary language and region subtags. For example, if a user requests
content in an implementation of Section 2.5 of [RFC3066] using the
language range "en-US", content labeled "en-Latn-US" will not match
the request. Therefore, it is important to know when script subtags
will customarily be used and when they ought not be used. In the
registry, the Suppress-Script field helps ensure greater
compatibility between the language tags generated according to the
rules in this document and language tags and tag processors or
consumers based on RFC 3066 by defining when users SHOULD NOT include
a script subtag with a particular primary language subtag.
Extended language subtags (type ’extlang’ in the registry; see
Section 3.1) also appear between the primary language and region
subtags and are reserved for future standardization. Applications
might benefit from their judicious use in forming language tags in
the future. Similar recommendations are expected to apply to their
use as apply to script subtags.
Standards, protocols, and applications that reference this document
normatively but apply different rules to the ones given in this
section MUST specify how the procedure varies from the one given
here.
The choice of subtags used to form a language tag SHOULD be guided by
the following rules:
1. Use as precise a tag as possible, but no more specific than is
justified. Avoid using subtags that are not important for
distinguishing content in an application.
* For example, ’de’ might suffice for tagging an email written
in German, while "de-CH-1996" is probably unnecessarily
precise for such a task.
2. The script subtag SHOULD NOT be used to form language tags unless
the script adds some distinguishing information to the tag. The
field ’Suppress-Script’ in the primary language record in the
registry indicates which script subtags do not add distinguishing
information for most applications.
* For example, the subtag ’Latn’ should not be used with the
primary language ’en’ because nearly all English documents are
written in the Latin script and it adds no distinguishing
information. However, if a document were written in English
mixing Latin script with another script such as Braille
(’Brai’), then it might be appropriate to choose to indicate
both scripts to aid in content selection, such as the
application of a style sheet.
3. If a tag or subtag has a ’Preferred-Value’ field in its registry
entry, then the value of that field SHOULD be used to form the
language tag in preference to the tag or subtag in which the
preferred value appears.
* For example, use ’he’ for Hebrew in preference to ’iw’.
4. The ’und’ (Undetermined) primary language subtag SHOULD NOT be
used to label content, even if the language is unknown. Omitting
the language tag altogether is preferred to using a tag with a
primary language subtag of ’und’. The ’und’ subtag MAY be useful
for protocols that require a language tag to be provided. The
’und’ subtag MAY also be useful when matching language tags in
certain situations.
5. The ’mul’ (Multiple) primary language subtag SHOULD NOT be used
whenever the protocol allows the separate tags for multiple
languages, as is the case for the Content-Language header in
HTTP. The ’mul’ subtag conveys little useful information:
content in multiple languages SHOULD individually tag the
languages where they appear or otherwise indicate the actual
language in preference to the ’mul’ subtag.
6. The same variant subtag SHOULD NOT be used more than once within
a language tag.
* For example, do not use "de-DE-1901-1901".
To ensure consistent backward compatibility, this document contains
several provisions to account for potential instability in the
standards used to define the subtags that make up language tags.
These provisions mean that no language tag created under the rules in
this document will become obsolete.
4.2. Meaning of the Language Tag
The relationship between the tag and the information it relates to is
defined by the context in which the tag appears. Accordingly, this
section gives only possible examples of its usage.
o For a single information object, the associated language tags
might be interpreted as the set of languages that is necessary for
a complete comprehension of the complete object. Example: Plain
text documents.
o For an aggregation of information objects, the associated language
tags could be taken as the set of languages used inside components
of that aggregation. Examples: Document stores and libraries.
o For information objects whose purpose is to provide alternatives,
the associated language tags could be regarded as a hint that the
content is provided in several languages and that one has to
inspect each of the alternatives in order to find its language or
languages. In this case, the presence of multiple tags might not
mean that one needs to be multi-lingual to get complete
understanding of the document. Example: MIME multipart/
alternative.
o In markup languages, such as HTML and XML, language information
can be added to each part of the document identified by the markup
structure (including the whole document itself). For example, one
could write <span lang="fr">C’est la vie.</span> inside a
Norwegian document; the Norwegian-speaking user could then access
a French-Norwegian dictionary to find out what the marked section
meant. If the user were listening to that document through a
speech synthesis interface, this formation could be used to signal
the synthesizer to appropriately apply French text-to-speech
pronunciation rules to that span of text, instead of applying the
inappropriate Norwegian rules.
Language tags are related when they contain a similar sequence of
subtags. For example, if a language tag B contains language tag A as
a prefix, then B is typically "narrower" or "more specific" than A.
Thus, "zh-Hant-TW" is more specific than "zh-Hant".
This relationship is not guaranteed in all cases: specifically,
languages that begin with the same sequence of subtags are NOT
guaranteed to be mutually intelligible, although they might be. For
example, the tag "az" shares a prefix with both "az-Latn"
(Azerbaijani written using the Latin script) and "az-Cyrl"
(Azerbaijani written using the Cyrillic script). A person fluent in
one script might not be able to read the other, even though the text
might be identical. Content tagged as "az" most probably is written
in just one script and thus might not be intelligible to a reader
familiar with the other script.
4.3. Length Considerations
[RFC3066] did not provide an upper limit on the size of language
tags. While RFC 3066 did define the semantics of particular subtags
in such a way that most language tags consisted of language and
region subtags with a combined total length of up to six characters,
larger registered tags were not only possible but were actually
registered.
Neither the language tag syntax nor other requirements in this
document impose a fixed upper limit on the number of subtags in a
language tag (and thus an upper bound on the size of a tag). The
language tag syntax suggests that, depending on the specific
language, more subtags (and thus a longer tag) are sometimes
necessary to completely identify the language for certain
applications; thus, it is possible to envision long or complex subtag
sequences.
4.3.1. Working with Limited Buffer Sizes
Some applications and protocols are forced to allocate fixed buffer
sizes or otherwise limit the length of a language tag. A conformant
implementation or specification MAY refuse to support the storage of
language tags that exceed a specified length. Any such limitation
SHOULD be clearly documented, and such documentation SHOULD include
what happens to longer tags (for example, whether an error value is
generated or the language tag is truncated). A protocol that allows
tags to be truncated at an arbitrary limit, without giving any
indication of what that limit is, has the potential for causing harm
by changing the meaning of tags in substantial ways.
In practice, most language tags do not require more than a few
subtags and will not approach reasonably sized buffer limitations;
see Section 4.1.
Some specifications or protocols have limits on tag length but do not
have a fixed length limitation. For example, [RFC2231] has no
explicit length limitation: the length available for the language tag
is constrained by the length of other header components (such as the
charset’s name) coupled with the 76-character limit in [RFC2047].
Thus, the "limit" might be 50 or more characters, but it could
potentially be quite small.
The considerations for assigning a buffer limit are:
Implementations SHOULD NOT truncate language tags unless the
meaning of the tag is purposefully being changed, or unless the
tag does not fit into a limited buffer size specified by a
protocol for storage or transmission.
Implementations SHOULD warn the user when a tag is truncated since
truncation changes the semantic meaning of the tag.
Implementations of protocols or specifications that are space
constrained but do not have a fixed limit SHOULD use the longest
possible tag in preference to truncation.
Protocols or specifications that specify limited buffer sizes for
language tags MUST allow for language tags of up to 33 characters.
Protocols or specifications that specify limited buffer sizes for
language tags SHOULD allow for language tags of at least 42
characters.
The following illustration shows how the 42-character recommendation
was derived. The combination of language and extended language
subtags was chosen for future compatibility. At up to 15 characters,
this combination is longer than the longest possible primary language
subtag (8 characters):
language = 3 (ISO 639-2; ISO 639-1 requires 2)
extlang1 = 4 (each subsequent subtag includes ’-’)
extlang2 = 4 (unlikely: needs prefix="language-extlang1")
extlang3 = 4 (extremely unlikely)
script = 5 (if not suppressed: see Section 4.1)
region = 4 (UN M.49; ISO 3166 requires 3)
variant1 = 9 (MUST have language as a prefix)
variant2 = 9 (MUST have language-variant1 as a prefix)
total = 42 characters
Figure 7: Derivation of the Limit on Tag Length
4.3.2. Truncation of Language Tags
Truncation of a language tag alters the meaning of the tag, and thus
SHOULD be avoided. However, truncation of language tags is sometimes
necessary due to limited buffer sizes. Such truncation MUST NOT
permit a subtag to be chopped off in the middle or the formation of
invalid tags (for example, one ending with the "-" character).
This means that applications or protocols that truncate tags MUST do
so by progressively removing subtags along with their preceding "-"
from the right side of the language tag until the tag is short enough
for the given buffer. If the resulting tag ends with a single-