TEI in practice

The TEI header

TEI Basics

A TEI document consists of two main sections: the <teiHeader> and <text>. These two sections are child elements of the <TEI> root element. While the <text> element contains the document text (the encoded poem, letter or other textual object), the <teiHeader> provides for an extensive metadata record of the object, both the original analogue object (if appropriate) as well as the encoded version. A schematic of the basic structure is below in figure 1.

1
2
3
4
5
6
7
8
<TEI xmlns="http://www.tei-c.org/ns/1.0">
     <teiHeader>
            <!-- metadata -->
     </teiHeader>
     <text>
             <!-- transcription -->
     </text>
</TEI>

Figure 1: a schematic of the basic structure of the TEI.

Structure of the TEI Header

The main function of the <teiHeader> is to provide bibliographic record about the electronic document. It includes four main sections (or child elements), not all of which are required for conformant TEI:

  • <fileDesc>; a bibliographic record of the electronic text as well as the source from which the electronic text is derived;
  • <encodingDesc>: documentation of the encoding and editorial principles used in tagging the electronic text;
  • <profileDesc>: terms for indexing, searching and retrieval;
  • <revisionDesc>: a record of changes made to the electronic document.

Figure 2. shows the order of these elements, if they are all used. Although the only required child element is <fileDesc>.

1
2
3
4
5
6
<teiHeader>
    <fileDesc></fileDesc>
    <encodingDesc></encodingDesc>
    <profileDesc></profileDesc>
    <revisionDesc></revisionDesc>
</teiHeader>
Figure 2. The child elements of <teiHeader> and their required order.

<fileDesc> in more detail

<fileDesc> is the only child element of the <teiHeader> that is mandatory in all TEI documents. <fileDesc>, in turn, must include three child elements to be conformant:  <titleStmt>, <publicationStmt> and <sourceDesc>.

  • <titleStmt> contains child elements that provides basic metadata about the document, including title of the resource, author and/or editor names, as well as the names and roles of other people who contributed to the creation of the electronic document;
  • <publicationStmt> contains basic child elements regarding publication information of the electronic text, including publlisher name and address, copyright information, and publication date;
  • <sourceDesc> contains child elements that describe the original source from which the electronic text was created. For instance, it may contain a detailed description of a manuscript or book. 

Figure 3. shows a basic TEI header structure with all mandatory elements:

1
2
3
4
5
6
7
8
9
10
11
12
13
<teiHeader>
    <fileDesc>
        <titleStmt>
             <title>Title</title>
        </titleStmt>
        <publicationStmt>
             <p>Publication Information</p>
        </publicationStmt>
        <sourceDesc>
            <p>Information about the source</p>
        </sourceDesc>
    </fileDesc>
</teiHeader>

Figure 3: Mandatory TEI elements.



Besides the three mandatory child elements of <fileDesc> described above, there are also four optional elements, the order of which are mandated by the Guidelines:

  • <editionStmt>: groups information relating to one edition of a text;
  • <extent>: describes the approximate size of a text stored on some carrier medium or of some other object, digital or non-digital, specified in any convenient units;
  • <seriesStmt>: groups information about the series, if any, to which a publication belongs;
  • <noteStmt>: collects together any notes providing information about a text additional to that recorded in other parts of the bibliographic description.

Options for encoding TEI header elements

As described above, many of the elements contained in the <teiHeader> contain child elements that provide further structuring of the content.  With many of these elements an editor can choose between a simple prose description or a more detailed structuring of the content using additional child elements. <encodingDesc> is a case in point. One option is to use a simple <p> tag for a prose description as in figure 4:

1
2
3
<encodingDesc>
    <p>Original spelling and typography is retained, except that long s and ligatured forms are not encoded.</p>
</encodingDesc>
Figure 4. A simple prose description of <encodingDesc>

However, the <encodingDesc> could  encoding to provide much more structure via child elements such as figure 5:

1
2
3
4
5
6
<encodingDesc>
   <projectDesc>describes the overall project purpose and process</projectDesc>
   <samplingDecl>documents rational for text sampling or selection in case of parts of the text or corpus have been omitted</samplingDecl>
   <editorialDecl>explains editorial principles of encoding or transcribing texts</editorialDecl>
   <charDecl>provides information about nonstandard characters and glyphs</charDecl>
</encodingDesc>
Figure 5. A detailed encoding of <encodingDesc>.

Both types of encoding are absolutely correct. The decision as to which one to use will depend on the goals of the project the the purpose of the encoded texts. For example, if it is important for users to be able to search on how nonstandard characters were handled in the encoding, then the second example would be more suitable as the use of the <charDecl> would allow that element to be isolated for search purposes.

Chapter Two of the TEI Guidelines explains the <teiHeader> and all the child elements that can be used. The exercises in this unit also go into the Header in more detail. 


Further reading

The TEI header. The TEI Guidelines. <http://www.tei-c.org/release/doc/tei-p5-doc/en/html/HD.html>

Module 2: The TEI header, TEI by Example. <http://teibyexample.org/modules/TBED02v00.htm>