Wiktionary data extraction errors and warnings

Inflection check

List of different kinds of inflection tables. When wiktextract parses word heads and tables, it assigns the forms it encounters with tags that describe grammatical or contextual information. The tags and forms that are found in head sections and tables are kept separate from other head section and table tags, and later they are merged with other heads and tables into table types that all contain the same number of word forms with the same tags for those forms.

The information presented here is mostly for debugging, but it can also be used to find interesting word paradigms and to hunt down mistakes, typoes and badly formated Wiktionary entries. A table type that has only a few unique instances is quite likely to contain some kind of minor error in the original data.

Language ⏶ Table forms Errors (% affected words) Language Table forms ⏷ Errors (% affected words)
갈리시아어 1 2 (100.00%) 일본어 80 132 (45.99%)
고전 만다어 1 2 (100.00%) 중국어 10 4 (1.22%)
따이담어 1 2 (100.00%) 한국어 8 24 (98.87%)
라디노어 1 2 (100.00%) 스와힐리어 5 10 (90.00%)
라오어 1 2 (100.00%) 러시아어 5 22 (88.14%)
라틴어 2 4 (100.00%) 아랍어 5 10 (13.16%)
러시아어 5 22 (88.14%) 베트남어 3 2 (0.27%)
마라티어 1 0 (0.00%) 모로코 아랍어 3 4 (5.88%)
마카오어 1 2 (100.00%) 영어 2 2 (99.71%)
모로코 아랍어 3 4 (5.88%) 라틴어 2 4 (100.00%)
버마어 2 2 (3.70%) 카라칼파크어 2 4 (66.67%)
베트남어 3 2 (0.27%) 조지아어 2 2 (83.72%)
벨라루스어 1 2 (100.00%) 카자흐어 2 2 (7.69%)
스와힐리어 5 10 (90.00%) 우크라이나어 2 0 (0.00%)
스코틀랜드 게일어 1 2 (100.00%) 버마어 2 2 (3.70%)
스페인어 1 2 (100.00%) 프랑스어 1 2 (100.00%)
슬로바키아어 1 2 (100.00%) 이탈리아어 1 2 (100.00%)
아랍어 5 10 (13.16%) 체코어 1 2 (100.00%)
아르메니아어 1 4 (100.00%) 스페인어 1 2 (100.00%)
아스투리아스어 1 2 (100.00%) 타타르어 1 2 (100.00%)
에스페란토 1 6 (100.00%) 에스페란토 1 6 (100.00%)
영어 2 2 (99.71%) 라디노어 1 2 (100.00%)
우드무르트어 1 2 (100.00%) 벨라루스어 1 2 (100.00%)
우크라이나어 2 0 (0.00%) 크림타타르어 1 2 (100.00%)
이탈리아어 1 2 (100.00%) 폴란드어 1 2 (100.00%)
일본어 80 132 (45.99%) 마카오어 1 2 (100.00%)
제주어 1 16 (100.00%) 갈리시아어 1 2 (100.00%)
조지아어 2 2 (83.72%) 슬로바키아어 1 2 (100.00%)
중국어 10 4 (1.22%) 카라임어 1 2 (100.00%)
체코어 1 2 (100.00%) 제주어 1 16 (100.00%)
츠와나어 1 2 (100.00%) 스코틀랜드 게일어 1 2 (100.00%)
카라임어 1 2 (100.00%) 츠와나어 1 2 (100.00%)
카라칼파크어 2 4 (66.67%) 우드무르트어 1 2 (100.00%)
카자흐어 2 2 (7.69%) 키르기스어 1 0 (0.00%)
콥트어 1 6 (100.00%) 아르메니아어 1 4 (100.00%)
크림타타르어 1 2 (100.00%) 크메르어 1 0 (0.00%)
크메르어 1 0 (0.00%) 콥트어 1 6 (100.00%)
키르기스어 1 0 (0.00%) 마라티어 1 0 (0.00%)
타타르어 1 2 (100.00%) 라오어 1 2 (100.00%)
티그레어 1 2 (100.00%) 따이담어 1 2 (100.00%)
폴란드어 1 2 (100.00%) 아스투리아스어 1 2 (100.00%)
프랑스어 1 2 (100.00%) 티그레어 1 2 (100.00%)
한국어 8 24 (98.87%) 고전 만다어 1 2 (100.00%)

This page is a part of the kaikki.org machine-readable dictionary. This dictionary is based on structured data extracted on 2026-10-02 from the kowiktionary dump dated 2026-10-01 using wiktextract (1a05e46 and e3d6d4e). The data shown on this site has been post-processed and various details (e.g., extra categories) removed, some information disambiguated, and additional data merged from other sources. See the raw data download page for the unprocessed wiktextract data.

If you use this data in academic research, please cite Tatu Ylonen: Wiktextract: Wiktionary as Machine-Readable Structured Data, Proceedings of the 13th Conference on Language Resources and Evaluation (LREC), pp. 1317-1325, Marseille, 20-25 June 2022. Linking to the relevant page(s) under https://kaikki.org would also be greatly appreciated.