ConSOLE 19 2011 University of Groningen

Multiple sequence alignment in historical linguistics. A sound class based approach

Johann-Mattis List

Heinrich Heine University Düsseldorf

listm@phil.uni-duesseldorf.de
Multiple Sequence AlignmentHistorical LinguisticsBulgarian dialects

Abstract

In this paper, a new method for multiple sequence alignment in historical linguistics is presented. The algorithm is based on the traditional framework of progressive multiple sequence alignment (cf. Durbin et al. 2002:143-149) whose shortcomings are further enhanced by (1) a sound class representation of phonetic sequences (cf. Dolgopolsky 1986, Turchin et al. 2010) accompanied by specific scoring functions, (2) the modification of gap scores based on prosodic context, (3) a new method for the detection of swapped sites in already aligned sequences. The algorithm is implemented as part of the LingPy library (http://lingulist.de/lingpy), a suite of open source Python modules for various tasks in quantitative historical linguistics. The method was tested on a benchmark dataset of 152 manually edited multiple alignments covering data for 192 Bulgarian dialects (Prokić et al. 2009). The results show that the new method yields alignments which differ only in 5 % of all sequences from the gold standard.

Access & Citation

Citation Formats

APA Style

Johann-Mattis List (2011). multiple sequence alignment in historical linguistics. a sound class based approach. In Proceedings of ConSOLE 19, edited by Enrico Boone, Kathrin Linke, Maartje Schulpen, (pp. 241-260).

BibTeX

@inproceedings{List-multiplesequence-2012, title={Multiple sequence alignment in historical linguistics. A sound class based approach}, author={Johann-Mattis List}, booktitle={Proceedings of ConSOLE 19}, year={2011}, pages={241-260}, editor={Enrico Boone and Kathrin Linke and Maartje Schulpen} }