Numeric character reference

A numeric character reference (NCR) is a common markup construct used in SGML and SGML-derived markup languages such as HTML and XML. It consists of a short sequence of characters that, in turn, represents a single character. Since WebSgml, XML and HTML 4, the code points of the Universal Character Set (UCS) of Unicode are used. NCRs are typically used in order to represent characters that are not directly encodable in a particular document (for example, because they are international characters that don't fit in the 8-bit character set being used, or because they have special syntactic meaning in the language). When the document is interpreted by a markup-aware reader, each NCR is treated as if it were the character it represents.

Examples

In SGML, HTML, and XML, the following are all valid numeric character references for the Greek capital letter Sigma

Numerical character reference of U+03A3 Σ GREEK CAPITAL LETTER SIGMA
(3A3₁₆ = 931)
Unicode character	Numerical base	Numerical reference in markup	Effect
U+03A3	Decimal	Σ	Σ
U+03A3	Decimal	Σ	Σ
U+03A3	Hexadecimal	Σ	Σ
U+03A3	Hexadecimal	Σ	Σ
U+03A3	Hexadecimal	Σ	Σ

In SGML, HTML, and XML, the following are all valid numeric character references for the Latin capital letter AE

Numerical character reference of U+00C6 Æ LATIN CAPITAL LETTER AE
Unicode character	Numerical base	Numerical reference in markup	Effect
U+00C6	Decimal	Æ	Æ
U+00C6	Hexadecimal	Æ	Æ

In SGML, HTML, and XML, the following are all valid numeric character references for the Latin small letter sharp s ß

Numerical character reference of U+00DF ß LATIN SMALL LETTER SHARP S
Unicode character	Numerical base	Numerical reference in markup	Effect
U+00DF	Decimal	ß	ß
U+00DF	Hexadecimal	ß	ß

List of numeric character references for the printable ASCII characters:

Unicode character	Character Reference (decimal)	Character Reference (hexadecimal)	Effect
U+0020			(space)
U+0021	!	!	!
U+0022	"	"	"
U+0023	#	#	#
U+0024	$	$	$
U+0025	%	%	%
U+0026	&	&	&
U+0027	'	'	'
U+0028	(	(	(
U+0029	)	)	)
U+002A	*	*	*
U+002B	+	+	+
U+002C	,	,	,
U+002D	-	-	-
U+002E	.	.	.
U+002F	/	/	/
U+0030	0	0	0
U+0031	1	1	1
U+0032	2	2	2
U+0033	3	3	3
U+0034	4	4	4
U+0035	5	5	5
U+0036	6	6	6
U+0037	7	7	7
U+0038	8	8	8
U+0039	9	9	9
U+003A	:	:	:
U+003B	;	;	;
U+003C	<	<	<
U+003D	=	=	=
U+003E	>	>	>
U+003F	?	?	?
U+0040	@	@	@
U+0041	A	A	A
U+0042	B	B	B
U+0043	C	C	C
U+0044	D	D	D
U+0045	E	E	E
U+0046	F	F	F
U+0047	G	G	G
U+0048	H	H	H
U+0049	I	I	I
U+004A	J	J	J
U+004B	K	K	K
U+004C	L	L	L
U+004D	M	M	M
U+004E	N	N	N
U+004F	O	O	O
U+0050	P	P	P
U+0051	Q	Q	Q
U+0052	R	R	R
U+0053	S	S	S
U+0054	T	T	T
U+0055	U	U	U
U+0056	V	V	V
U+0057	W	W	W
U+0058	X	X	X
U+0059	Y	Y	Y
U+005A	Z	Z	Z
U+005B	[	[	[
U+005C	\	\	\
U+005D	]	]	]
U+005E	^	^	^
U+005F	_	_	_
U+0060	`	`	`
U+0061	a	a	a
U+0062	b	b	b
U+0063	c	c	c
U+0064	d	d	d
U+0065	e	e	e
U+0066	f	f	f
U+0067	g	g	g
U+0068	h	h	h
U+0069	i	i	i
U+006A	j	j	j
U+006B	k	k	k
U+006C	l	l	l
U+006D	m	m	m
U+006E	n	n	n
U+006F	o	o	o
U+0070	p	p	p
U+0071	q	q	q
U+0072	r	r	r
U+0073	s	s	s
U+0074	t	t	t
U+0075	u	u	u
U+0076	v	v	v
U+0077	w	w	w
U+0078	x	x	x
U+0079	y	y	y
U+007A	z	z	z
U+007B	{	{	{
U+007C	\|	\|	\|
U+007D	}	}	}
U+007E	~	~	~

Discussion

Markup languages are typically defined in terms of UCS or Unicode characters. That is, a document consists, at its most fundamental level of abstraction, of a sequence of characters, which are abstract units that exist independently of any encoding.

Ideally, when the characters of a document utilizing a markup language are encoded for storage or transmission over a network as a sequence of bits, the encoding that is used will be one that supports representing each and every character in the document, if not in the whole of Unicode, directly as a particular bit sequence.

Sometimes, though, for reasons of convenience or due to technical limitations, documents are encoded with an encoding that cannot represent some characters directly. For example, the widely used encodings based on ISO 8859 can only represent, at most, 256 unique characters as one 8-bit byte each.

Documents are rarely, in practice, ever allowed to use more than one encoding internally, so the onus is usually on the markup language to provide a means for document authors to express unencodable characters in terms of encodable ones. This is generally done through some kind of "escaping" mechanism.

The SGML-based markup languages allow document authors to use special sequences of characters from the ASCII range (the first 128 code points of Unicode) to represent, or reference, any Unicode character, regardless of whether the character being represented is directly available in the document's encoding. These special sequences are character references.

Character references that are based on the referenced character's UCS or Unicode code point are called numeric character references. In HTML 4 and in all versions of XHTML and XML, the code point can be expressed either as a decimal (base 10) number or as a hexadecimal (base 16) number. The syntax is as follows:

Character U+0026 (ampersand), followed by character U+0023 (number sign), followed by one of the following choices:

one or more decimal digits zero (U+0030) through nine (U+0039); or
character U+0078 ("x") followed by one or more hexadecimal digits, which are zero (U+0030) through nine (U+0039), Latin capital letter A (U+0041) through F (U+0046), and Latin small letter a (U+0061) through f (U+0066);

all followed by character U+003B (semicolon). Older versions of HTML disallowed the hexadecimal syntax.

The characters that comprise a numeric character reference can be represented in every character encoding used in computing and telecommunications today, so there is no risk of the reference itself being unencodable.

There is another kind of character reference called a character entity reference, which allows a character to be referred to by a name instead of a number. (Naming a character creates a character entity.) HTML defines some character entities, but not many; all other characters can only be included by direct encoding or using NCRs.

Compatibility issues

In the initial versions of SGML and HTML, numeric character references were interpreted in relationship to the document character encoding, rather than Unicode. For Latin-script documents, numeric character references to characters between x80 and x9F in those documents will not be correct against Unicode, and must be recoded. HTML standards prior to HTML 4 only supported Western Latin script documents: the treatment of character references above #7F may vary between applications and national conventions.

For example, as mentioned above, the correct numeric character reference for the Euro sign "€" U+20AC when using Unicode is decimal € and hexadecimal €. However, if using tools supporting obsolete implementations of HTML, the reference  (Euro in Cp1252 code page) or ¤ (Euro in ISO/IEC 8859-15 ) may work.

As another example, if some text was created originally MacRoman character set, the left double quotation mark “ will be represented with codepoint xD2. This will not display properly in a system expecting a document encoded as UTF-8, ISO 8859-1, or CP1252, where this code point is occupied by the letter Ò. The correct numeric character reference for “ in HTML 4 and newer is “, because U+201C is its UCS code. In some systems, the named character reference “ may also be available.

]

References