decision is one of DNS root-level namespace policy and hence falls
to ICANN although we would expect ICANN to pay careful attention
to any technical, operational, or security recommendations that
may be produced by other bodies.
o Finally, if IDN labels are to be placed in the root zone, there
are issues associated with how they are to be encoded and
deployed. This area may have implications for work that has been
done, or should be done, in the IETF.
5. Specific Recommendations for Next Steps
Consistent with the framework described above, the IAB offers these
recommendations as steps for further consideration in the identified
groups.
5.1. Reduction of Permitted Character List
Generalize from the original "hostname" rules to non-ASCII
characters, permitting as few characters as possible to do that job.
This would involve a restrictive model for characters permitted in
IDN labels, thus contrasting with the approach used to develop the
original IDNA/Nameprep tables. That approach was to include all
Unicode characters that there was not a clear reason to exclude.
The specific recommendation here is to specify such internationalized
hostnames. Such an activity would fall to the IETF, although the
task of developing the appropriate list of permitted characters will
require effort both in the IETF and elsewhere. The effort should be
as linguistically and culturally sensitive as possible, but smooth
and effective operation of the DNS, including minimizing of
complexity, should be primary goals. The following should be
considered as possible mechanisms for achieving an appropriate
minimum number of characters.
5.1.1. Elimination of All Non-Language Characters
Unicode characters that are not needed to write words or numbers in
any of the world’s languages should be eliminated from the list of
characters that are appropriate in DNS labels. In addition to such
characters as those used for box-drawing and sentence punctuation,
this should exclude punctuation for word structure and other
delimiters. While DNS labels may conveniently be used to express
words in many circumstances, the goal is not to express words (or
sentences or phrases), but to permit the creation of unambiguous
labels with good mnemonic value.
5.1.2. Elimination of Word-Separation Punctuation
The inclusion of the hyphen in the original hostname rules is a
historical artifact from an older, flat, namespace. The community
should consider whether it is appropriate to treat it as a simple
legacy property of ASCII names and not attempt to generalize it to
other scripts. We might, for example, not permit claimed equivalents
to the hyphen from other scripts to be used in IDNs. We might even
consider banning use of the hyphen itself in non-ASCII strings or,
less restrictively, strings that contained non-Latin characters.
5.2. Updating to New Versions of Unicode
As new scripts, to support new languages, continue to be added to
Unicode, it is important that IDNA track updates. If it does not do
so, but remains "stuck" at 3.2 or some single later version, it will
not be possible to include labels in the DNS that are derived from
words in languages that require characters that are available only in
later versions. Making those upgrades is difficult, and will
continue to be difficult, as long as new versions require, not just
addition of characters, but changes to canonicalization conventions,
normalization tables, or matching procedures (see Section 3.1).
Anything that can be done to lower complexity and simplify forward
transitions should be seriously considered.
5.3. Role and Uses of the DNS
We wish to remind the community that there are boundaries to the
appropriate uses of the DNS. It was designed and implemented to
serve some specific purposes. There are additional things that it
does well, other things that it does badly, and still other things it
cannot do at all. No amount of protocol work on IDNs will solve
problems with alternate spellings, near-matches, searching for
appropriate names, and so on. Registration restrictions and
carefully-designed user interfaces can be used to reduce the risk and
pain of attempts to do some of these things gone wrong, as well as
reducing the risks of various sort of deliberate bad behavior, but,
beyond a certain point, use of the DNS simply because it is available
becomes a bad tradeoff. The tradeoff may be particularly unfortunate
when the use of IDNs does not actually solve the proposed problem.
For example, internationalization of DNS names does not eliminate the
ASCII protocol identifiers and structure of URIs [RFC3986] and even
IRIs [RFC3987]. Hence, DNS internationalization itself, at any or
all levels of the DNS tree, is not a sufficient response to the
desire of populations to use the Internet entirely in their own
languages and the characters associated with those languages.
These issues are discussed at more length, and alternatives
presented, in [RFC2825], [RFC3467], [INDNS], and [DNS-Choices].
5.4. Databases of Registered Names
In addition to their presence in the DNS, IDNs introduce issues in
other contexts in which domain names are used. In particular, the
design and content of databases that bind registered names to
information about the registrant (commonly described as "whois"
databases) will require review and updating. For example, the whois
protocol itself [RFC3912] has no standard capability for handling
non-ASCII text: one cannot search consistently for, or report, either
a DNS name or contact information that is not in ASCII characters.
This may provide some additional impetus for a switch to IRIS
[RFC3981] [RFC3982] but also raises a number of other questions about
what information, and in what languages and scripts, should be
included or permitted in such databases.
6. Security Considerations
This document is simply a discussion of IDNs and IDNA issues; it
raises no new security concerns. However, if some of its
recommendations to reduce IDNA complexity, the number of available
characters, and various approaches to constraining the use of
confusable characters, are followed and prove successful, the risks
of name spoofing and other problems may be reduced.
7. Acknowledgements
The contributions to this report from members of the IAB-IDN ad hoc
committee are gratefully acknowledged. Of course, not all of the
members of that group endorse every comment and suggestion of this
report. In particular, this report does not claim to reflect the
views of the Unicode Consortium as a whole or those of particular
participants in the work of that Consortium.
The members of the ad hoc committee were: Rob Austein, Leslie Daigle,
Tina Dam, Mark Davis, Patrik Faltstrom, Scott Hollenbeck, Cary Karp,
John Klensin, Gervase Markham, David Meyer, Thomas Narten, Michael
Suignard, Sam Weiler, Bert Wijnen, Kurt Zeilenga, and Lixia Zhang.
Thanks are due to Tina Dam and others associated with the ICANN IDN
Working Group for contributions of considerable specific text, to
Marcos Sanz and Paul Hoffman for careful late-stage reading and
extensive comments, and to Pete Resnick for many contributions and
comments, both in conjunction with his former IAB service and
subsequently. Olaf M. Kolkman took over IAB leadership for this
document after Patrik Faltstrom and Pete Resnick stepped down in
March 2006.
Members of the IAB at the time of approval of this document were:
Bernard Aboba, Loa Andersson, Brian Carpenter, Leslie Daigle, Patrik
Faltstrom, Bob Hinden, Kurtis Lindqvist, David Meyer, Pekka Nikander,
Eric Rescorla, Pete Resnick, Jonathan Rosenberg and Lixia Zhang.
8. References
8.1. Normative References
[ISO10646] International Organization for Standardization,
"Information Technology - Universal Multiple-
Octet Coded Character Set (UCS) - Part 1:
Architecture and Basic Multilingual Plane"",
ISO/IEC 10646-1:2000, October 2000.
[RFC3454] Hoffman, P. and M. Blanchet, "Preparation of
Internationalized Strings ("stringprep")",
RFC 3454, December 2002.
[RFC3490] Faltstrom, P., Hoffman, P., and A. Costello,
"Internationalizing Domain Names in Applications
(IDNA)", RFC 3490, March 2003.
[RFC3491] Hoffman, P. and M. Blanchet, "Nameprep: A
Stringprep Profile for Internationalized Domain
Names (IDN)", RFC 3491, March 2003.
[RFC3492] Costello, A., "Punycode: A Bootstring encoding of
Unicode for Internationalized Domain Names in
Applications (IDNA)", RFC 3492, March 2003.
[Unicode32] The Unicode Consortium, "The Unicode Standard,
Version 3.0", 2000.
(Reading, MA, Addison-Wesley, 2000. ISBN
0-201-61633-5). Version 3.2 consists of the
definition in that book as amended by the Unicode
Standard Annex #27: Unicode 3.1
(http://www.unicode.org/reports/tr27/) and by the
Unicode Standard Annex #28: Unicode 3.2
(http://www.unicode.org/reports/tr28/).
8.2. Informative References
[DNS-Choices] Faltstrom, P., "Design Choices When Expanding
DNS", Work in Progress, June 2005.
[ICANNv1] ICANN, "Guidelines for the Implementation of
Internationalized Domain Names, Version 1.0",
March 2003, <http://www.icann.org/general/
idn-guidelines-20jun03.htm>.
[ICANNv2] ICANN, "Guidelines for the Implementation of
Internationalized Domain Names, Version 2.0",
November 2005, <http://www.icann.org/general/
idn-guidelines-20sep05.htm>.
[IESG-IDN] Internet Engineering Steering Group (IESG), "IESG
Statement on IDN", IESG Statements IDN Statement,
February 2003, <http://www.ietf.org/IESG/
STATEMENTS/IDNstatement.txt>.
[INDNS] National Research Council, "Signposts in
Cyberspace: The Domain Name System and Internet
Navigation", National Academy Press ISBN 0309-
09640-5 (Book) 0309-54979-5 (PDF), 2005, <http://
www7.nationalacademies.org/cstb/pub_dns.html>.
[ISO.2022.1986] International Organization for Standardization,
"Information Processing: ISO 7-bit and 8-bit
coded character sets: Code extension techniques",
ISO Standard 2022, 1986.
[ISO.646.1991] International Organization for Standardization,
"Information technology - ISO 7-bit coded
character set for information interchange",
ISO Standard 646, 1991.
[ISO.8859.2003] International Organization for Standardization,
"Information processing - 8-bit single-byte coded
graphic character sets - Part 1: Latin alphabet
No. 1 (1998) - Part 2: Latin alphabet No. 2
(1999) - Part 3: Latin alphabet No. 3 (1999) -
Part 4: Latin alphabet No. 4 (1998) - Part 5:
Latin/Cyrillic alphabet (1999) - Part 6: Latin/
Arabic alphabet (1999) - Part 7: Latin/Greek
alphabet (2003) - Part 8: Latin/Hebrew alphabet
(1999) - Part 9: Latin alphabet No. 5 (1999) -
Part 10: Latin alphabet No. 6 (1998) - Part 11:
Latin/Thai alphabet (2001) - Part 13: Latin
alphabet No. 7 (1998) - Part 14: Latin alphabet
No. 8 (Celtic) (1998) - Part 15: Latin alphabet
No. 9 (1999) - Part 16: Part 16: Latin alphabet
No. 10 (2001)", ISO Standard 8859, 2003.
[RFC2277] Alvestrand, H., "IETF Policy on Character Sets
and Languages", BCP 18, RFC 2277, January 1998.
[RFC2825] IAB and L. Daigle, "A Tangled Web: Issues of
I18N, Domain Names, and the Other Internet
protocols", RFC 2825, May 2000.
[RFC3066] Alvestrand, H., "Tags for the Identification of
Languages", BCP 47, RFC 3066, January 2001.
[RFC3467] Klensin, J., "Role of the Domain Name System
(DNS)", RFC 3467, February 2003.
[RFC3536] Hoffman, P., "Terminology Used in
Internationalization in the IETF", RFC 3536,
May 2003.
[RFC3743] Konishi, K., Huang, K., Qian, H., and Y. Ko,
"Joint Engineering Team (JET) Guidelines for
Internationalized Domain Names (IDN) Registration
and Administration for Chinese, Japanese, and
Korean", RFC 3743, April 2004.
[RFC3912] Daigle, L., "WHOIS Protocol Specification",
RFC 3912, September 2004.
[RFC3981] Newton, A. and M. Sanz, "IRIS: The Internet
Registry Information Service (IRIS) Core
Protocol", RFC 3981, January 2005.
[RFC3982] Newton, A. and M. Sanz, "IRIS: A Domain Registry
(dreg) Type for the Internet Registry Information
Service (IRIS)", RFC 3982, January 2005.
[RFC3986] Berners-Lee, T., Fielding, R., and L. Masinter,
"Uniform Resource Identifier (URI): Generic
Syntax", STD 66, RFC 3986, January 2005.
[RFC3987] Duerst, M. and M. Suignard, "Internationalized
Resource Identifiers (IRIs)", RFC 3987,
January 2005.
[RFC4185] Klensin, J., "National and Local Characters for
DNS Top Level Domain (TLD) Names", RFC 4185,
October 2005.
[RFC4290] Klensin, J., "Suggested Practices for
Registration of Internationalized Domain Names
(IDN)", RFC 4290, December 2005.
[RFC4645] Ewell, D., "Initial Language Subtag Registry",
RFC 4645, September 2006.
[RFC4646] Phillips, A. and M. Davis, "Tags for Identifying
Languages", BCP 47, RFC 4646, September 2006.
[UTR] Unicode Consortium, "Unicode Technical Reports",
<http://www.unicode.org/reports/>.
[UTR36] Davis, M. and M. Suignard, "Unicode Technical
Report #36: Unicode Security Considerations",
November 2005, <http://www.unicode.org/draft/
reports/tr36/tr36.html>.
[UTR39] Davis, M. and M. Suignard, "Unicode Technical
Standard #39 (proposed): Unicode Security
Considerations", July 2005, <http://
www.unicode.org/draft/reports/tr39/tr39.html>.
[Unicode-PR29] The Unicode Consortium, "Public Review Issue #29:
Normalization Issue", Unicode PR 29,
February 2004.
[Unicode10] The Unicode Consortium, "The Unicode Standard,
Version 1.0", 1991.
[W3C-Localization] Ishida, R. and S. Miller, "Localization vs.
Internationalization", W3C International/
questions/qa-i18n.txt, December 2005.
[net-utf8] Klensin, J. and M. Padlipsky, "Unicode Format for
Network Interchange", Work in Progress,
April 2006.
Authors’ Addresses
John C Klensin
1770 Massachusetts Ave, #322
Cambridge, MA 02140
USA
Phone: +1 617 491 5735
EMail: john-ietf@jck.com
Patrik Faltstrom
Cisco Systems
EMail: paf@cisco.com
Cary Karp
Swedish Museum of Natural History
Box 50007
Stockholm SE-10405
Sweden
Phone: +46 8 5195 4055
EMail: ck@nrm.museum
IAB
EMail: iab@iab.org
Full Copyright Statement
Copyright (C) The Internet Society (2006).
This document is subject to the rights, licenses and restrictions
contained in BCP 78, and except as set forth therein, the authors
retain all their rights.
This document and the information contained herein are provided on an
"AS IS" basis and THE CONTRIBUTOR, THE ORGANIZATION HE/SHE REPRESENTS
OR IS SPONSORED BY (IF ANY), THE INTERNET SOCIETY AND THE INTERNET
ENGINEERING TASK FORCE DISCLAIM ALL WARRANTIES, EXPRESS OR IMPLIED,
INCLUDING BUT NOT LIMITED TO ANY WARRANTY THAT THE USE OF THE
INFORMATION HEREIN WILL NOT INFRINGE ANY RIGHTS OR ANY IMPLIED
WARRANTIES OF MERCHANTABILITY OR FITNESS FOR A PARTICULAR PURPOSE.
Intellectual Property
The IETF takes no position regarding the validity or scope of any
Intellectual Property Rights or other rights that might be claimed to
pertain to the implementation or use of the technology described in
this document or the extent to which any license under such rights
might or might not be available; nor does it represent that it has
made any independent effort to identify any such rights. Information
on the procedures with respect to rights in RFC documents can be
found in BCP 78 and BCP 79.
Copies of IPR disclosures made to the IETF Secretariat and any
assurances of licenses to be made available, or the result of an
attempt made to obtain a general license or permission for the use of
such proprietary rights by implementers or users of this
specification can be obtained from the IETF on-line IPR repository at
http://www.ietf.org/ipr.
The IETF invites any interested party to bring to its attention any
copyrights, patents or patent applications, or other proprietary
rights that may cover technology that may be required to implement
this standard. Please address the information to the IETF at
ietf-ipr@ietf.org.
Acknowledgement
Funding for the RFC Editor function is provided by the IETF
Administrative Support Activity (IASA).