Basic XML rules: well-formed and valid XML

Building-blocks of XML documents

Elements

Within an XML document data is contained within elements. Elements have a start-tag and an end-tag. Start- and end-tags consist of an element name, which is a string of text such as ‘PersonName’, and a delimiter indicating the beginning and end of a tag. XML tags are delimited from the stored data by use of angle brackets (or inequality signs). The bracket < is used to indicate the beginning of an XML start-tag and > is used to indicate the end of a start-tag.

The end-tag starts with the delimiter </ and ends with the angle bracket >. The following would be an example of an XML element describing a person name. The content of the element, or in XML parlance, PCDATA (Parsed Character Data)  'Clark Kent' appears between the start- and end-tags:

1
<PersonName>Clark Kent</PersonName>

A special case are so-called empty elements. Empty elements do not contain PCDATA, and one tag functions as both the start and end tag. Empty elements are typically used as milestones. The XHTML <br /> which indicates line breaks would be typical of this.  This structure is frequently necessary to  indicate where breaks occur in a source text (e.g. a page or section break in a book or an article).

In the example of our element <PersonName> the element name is a string of characters. The rules for XML element names are that names must start with a letter or underscore, they can contain letters, digits, hyphens, underscores, and periods and they are case-sensitive. The following table shows a  examples of valid XML names:

Correct XML element names

<_newElement> </newElement>

Element names must start with a letter or underscore

<newElement> </newElement>

Element names are case-sensitive, start and end tag have to match

<my.new_Element-1></my.new_Element-1>

Element names can contain letters, digits, hyphens, underscores, and periods



Incorrect construction of element names leads to XML data that is not Well Formed. An XML processor will not process the data, but instead indicate an error message. The following table shows examples of illegal XML element names:

Illegal XML element names

<1Element> </1Element>

Element names cannot start with a digit, hyphen or period

< Element /> </ Element>

Element names cannot start with an initial space

<newElement> </Newelement>

Element names are case-sensitive and start and end tag must match

<xmlElement></xmlElement>

Element names cannot start with the string xml, XML, Xml, etc

<new Element> </new Element>

Element names cannot contain whitespace

Attributes

Another  building block of XML documents are attributes. Attributes are a way to add additional information in the form of name-value pairs to an XML element. Attributes modify, refine, or further delineate elements. If elements are thought of as nouns, attributes can be likened to adjectives. Attributes have to be placed after the element name of the start tag and they are separated by a whitespace from the element name. In the following example the attributes @first-name and @last-name with values are added to the XML element person:

<person first-name=”Henry” last-name=”James” />


Attributes can also be used to store data and sometimes it is difficult to decide if data should be stored as an attribute value or within an element. A clear benefit for storing information within an element is that data can be structured further by nesting other XML elements or with attributes. This cannot be done with an attribute value. The benefit of attributes is that they can be useful to describe an element and its content further and with schemas ,the data that is stored as attribute values can be restricted to specific data types and values.

For attributes, there are similar naming rules to elements. Names must start with a letter or underscore, they can contain letters, digits, hyphens, underscores, and periods, but cannot contain whitespace. Attribute values have to be quoted, they can contain alphanumeric characters, whitespace and various other characters such as period, hyphen, underscore, comma, etc. However, you have to be careful using single and double quotes. If single quotes are used as attribute value delimiter, they are not allowed in the value string. If double quotes are used as attribute value delimiter, they are not allowed in the value string. The following table contains examples of valid use of attributes.

Valid use of attributes in XML

<newElement attribute1=”attribute value: 1” />

Attribute name may contain digits, but cannot start with a digit. Attribute values may contain whitespaces, punctuation and alphanumeric characters in any order.

<person name=Rob Miller’ />

<person name=”Rob Miller” />

Single quotes or double quotes can be used as delimiter of attribute values.

<address owner=”Mary’s address” />

Single quotes can be used within a value string, but only if they are not deliminators.

<sentence spoken=’He said: “go”!’ />

Double quotes can be used within a value string, but only if they are not deliminators.


Incorrect attribute syntax leads to XML data that is not Well-Formed. An XML processor will not process the data, but instead indicate that there is an error. The following table shows examples of illegal attribute names and values:

Illegal use of attribute names and values in XML

<person name=Robert />

Attribute values must be quoted.

<person name=”Robert’ />

No mismatch between the quote delimiters.

<sentence spoken=”He said: “go”!” />

If double quotes are used as deliminators of a value string, they cannot be used in the value string itself.

<address owner=’Mary‘s address’ />

If single quotes are used as deliminators of a value string, they cannot be used in the value string itself.

<person first name=”Frank” />

An attribute name cannot contain whitespace characters.

<person 1stname=”Robert” />

Attribute names must start with a letter or underscore

<person name=”Robert” name=”James” />

An element cannot have multiple attributes with the same name.


Elements and attributes are the main building blocks of XML and they are essential for structuring  and modelling data. Other important components that are present in most XML documents are Processing instructions and XML and Unicode entities.


Processing instructions

Processing instructions contain instructions for an XML processor specifying what version of XML is used, what character encoding and/or what schema should be used to validate the XML document. Processing instructions can be easily recognised because they are written in the first lines of an XML document. Processing instruction do not have a start- and end-tags. However, they do have attributes to store information. For instance, the following processing instruction is the XML declaration and should be at the beginning of every XML document. It declares that this document is encoded according the XML standard using XML version 1.0 and the character encoding  UTF-8:

1
<?xml version="1.0" encoding="UTF-8"?>

Unicode entities

The text stored in an XML element is usually PCDATA (Parsed Character Data). This is a data definition used in XML documents and specifies that the five characters that are used to distinguish mark-up from data, such as angle brackets (for elements), single and double quotes (for attributes) and ampersand (for named entities), have to be escaped using Entity references. Entity references begin with & and end with a semicolon. The following list shows what Entities have to be used instead of the five illegal characters:

Character

XML entity

< 

&lt;

> 

&gt;

&

&amp;

'

&apos;

"

&quote;

 

Another form of Entity references are Character references for characters or symbols not contained in the ASCII character set. For instance, Unicode character references can be used within an XML document. Such character references start with &# and end with a semicolon and are directly embedded into an XML document. For instance, the Greek Capital Letter Pi has the character reference: &#x03A0;