Starting with TEI

TEI: in Medias Res

In Medias Res

We'll start this lesson by converting one part of the XML-encoded entry for lexicographer from Johnson's dictionary into TEI.

This is just to give you a taste of all those wonderful things to come and to give you a little more practice in using oXygen to rename existing elements and attributes or to create new ones. We'll cover different TEI elements in more depth later on.

  1. Select File > New from the menu
  2. Type "tei" in the search field
  3. Double-click on the template called All in the TEI5 folder
  4. Find the element <title> (which is the child of <titleStmt>, which itself is a child of <fileDesc>, which is a child of <teiHeader>) and give your file an appropriate tile by writing it between the <title> and </title>
  5. Add element after and write the name of Samuel Johnson
  6. Go the extra mile and add the element <respStmt> after </author>, and within it two more elements: <resp> and <name>. As the text value of describe your role in the production of this file, and in state your own name. Since this is just an exercise, you can call yourself whatever you want. Batman. Superman. Queen Victoria.
  7. Click on the downward pointing arrow next to <teiHeader> to fold the contents of the teiHeader.


  8. Go back to your previously saved XML file and copy everything starting with <entry> and ending with </grammar>; in your current TEI file, delete <p></p> inside <body></body> and then paste the copied contents from the old file instead.


  9. if you hover with your mouse over the elements <lemma> and <grammar> you will learn that they are "not allowed anywhere". That means that TEI doesn't know elements with such names. We'll fix them in due course. But there is another error in this document that you should be able to fix without any knowledge of TEI. Can you think what it is?
  10. Hover with your mouse over the underlined end tag to see what error oXygen will report there.


  11. When we asked you to copy and paste from your old file, we asked you to copy a fragment of xml which itself was not well-formed: <entry><lemma normalized='lexicographer'>Lexico‘grapher</lemma>. <grammar>n.s.</grammar> is not well-formed because each starting tag in XML must be matched by its end tag. That's why you need to enter </entry>in the empty row above </body>.
  12. Now, double-click the element name in <lemma>.

  13. Change the element name to <form>.


  14. Select @normalized and its value, then delete both.


  15. Add attribute @type to <form> and check out what options are suggested for values.

  16. Select attribute value lemma for @type.

  17. Select the text value of <form>.


  18. Control-click or right-click the selected word to bring up a contextual menu, then select Refactoring > Surround with Tags.

  19. Select orth from the dropdown by scrolling to it, or for a more efficient solution, start typing o in the field "Specify the tag".

  20. Add @norm to <orth>.

  21. Type in 'lexicographer' as value of @norm.

  22. Remember the "Fromat and Indent" icon that we introduced in the last unit? Find it and click on it.

  23. Finally, replace <grammar>n.s.</grammar> with <gramGrp><pos>n.s.</pos></gramGrp>.

  24. Make sure you have saved this file on your computer so that you can get back to it later on.

Congratulations! Your dictionary entry is beginning to look like it's speaking TEI! If you have followed this exercise, you will have noticed several things:

  • your <form> and <orth> elements have been accepted as proper TEI because oXygen has not underlined them in red.
  • we still have quite a bit of work to do on the rest of the entry: every element that's underlined in red indicates an error
  • because you selected a TEI template when you started working on this file, oXygen has been trying to help you as much as it can in suggesting possible elements and even possible attribute values
  • the reason why oXygen can do all these things has nothing to do with magic (unfortunately!), but with the very regimented world of XML schemas. We've talked briefly about them in Unit 2. A schema is a set of rules that the XML parser uses to validate your document against. The TEI schema tells oXygen: <lemma> is not a valid TEI element, but <form type="lemma"> is.
  • your TEI document is not both well-formed and valid. Make sure you understand the difference between the two concepts and, if necessary, revisit the corresponding section in Lesson 4 in Unit 2.

This completes your baptism by fire in things TEI. On the next page you will learn more about TEI itself and how to discover all those semi-exotic elements that populate the TEI universe.