SARIT and GRETIL Corpus Search

How to search

Type a word or phrase and press Search. The search runs in your browser over the TEI files of this repository: the SARIT texts and the GRETIL texts in GRETIL-corpustei/.

Scripts and transliteration

You can type in IAST (dharmakṣetre), Devanagari (धर्मक्षेत्रे) or plain ASCII. It does not matter which script the text is encoded in: before searching, both the texts and your query are converted to one form, which is lower-case IAST with the following unifications:

In the text or query is treated as
Devanagari IAST (क्ष → kṣa, ं → ṃ, ः → ḥ, ऽ → ‘)
ISO 15919 ṁ, r̥, l̥, ē, ō ṃ, ṛ, ḷ, e, o
/, । and //, ॥ the same single and double daṇḍa
’, ‘, ऽ (avagraha) '
Devanagari digits 0–9
Vedic accents, upper case ignored

So dharmakṣetre, धर्मक्षेत्रे and DHARMAKṢETRE find the same passages. Under More options you can also type in Harvard-Kyoto, Velthuis, SLP1 or ITRANS. The row of buttons under the search box inserts IAST letters.

Match modes

Ignoring things

The options can be combined. Whole words requires the match to begin and end at word boundaries (it cannot be combined with ignoring word division). Notes & variants also searches editorial notes and apparatus readings; the hits are labelled note or variant.

References and citation

For every hit the site shows as much as the TEI markup provides:

cite copies a full citation to the clipboard, including the source edition named in the TEI header and the file’s GitHub address. Download results saves all hits as a tab-separated file that opens in Excel or LibreOffice. context opens the passage in the reader, where you can page through the text, switch script and jump to a reference.

The search address in your browser’s location bar records the query and all options, so it can be bookmarked or cited.

Speed

The first search over the whole corpus (about 240 million characters) downloads the texts, which takes a little while; your browser then caches them and later searches are much faster. Restricting the search to SARIT, GRETIL or a few selected texts is quicker. Stop ends a search early and keeps the results found so far.

About the data

The site is rebuilt automatically whenever the repository changes: a script (website/_tools/extract.py) reads every TEI file, splits it into passages at the level of <p>, <lg>, <ab>, <head>, <item> etc., and records the references described above. Notes, <rdg> readings and similar material are kept apart from the main text. Where <choice> offers alternatives, the corrected/regularised/expanded form is searched.