Linguística Computacional e línguas indígenas brasileiras

Authors

DOI:

https://doi.org/10.47456/mmjhnb04

Keywords:

Linguística Computacional. Línguas indígenas brasileiras. Diversidade linguística. Recursos tecnológicos. Tipologia linguística.

Abstract

We present the challenges of Computational Linguistics in processing Brazilian Indigenous languages, addressing recent advancements, the main difficulties, and the perspectives on using available technologies in the preservation and revitalization of Indigenous languages. To this end, we provide a brief introduction to the field of Computational Linguistics and a review, based on Mager et al. (2018), of the most promising tools for research. Brazilian Indigenous languages represent a diverse linguistic set, with more than 250 languages spoken by different peoples. This diversity challenges conventional approaches in Computational Linguistics, which have struggled to address the particularities of these languages, as seen in the disconnect between technological solutions and their application to Indigenous languages. The lack of investment and social and political interest in research on native languages adds to the linguistic reality. The oral tradition and the absence of orthographic systems, as well as grammatical structures distinct from those of the more widely spoken languages, complicate the creation of databases necessary for training computational models and demand methodological innovations that consider their specificities. Furthermore, the adoption of Eurocentric views and concepts exacerbates the challenges. Through this exposition, we aim to contribute to the understanding of the current situation between Computational Linguistics and Brazilian Indigenous languages and emphasize the need for closer integration between research on technologies and native languages.

Author Biographies

  • Arthur Scandelari, University of Brasília

    Doutorando em Linguística (UnB) e bolsista da Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (Capes). Mestre em Linguística (UnB) e bacharel em Letras - Língua Portuguesa e Respectiva Literatura (UnB). Estudante dos grupos de pesquisa "Núcleo de Tipologia e Línguas Indígenas - NTL" (CNPq) e "Funcionalismo, Tipologia e Ensino" (CNPq). Pós-graduado em Direito Internacional pela Pontifícia Universidade Católica do Paraná (PUC-PR) e bacharel em Ciências Econômicas pela Universidade Federal do Paraná (UFPR).

  • Dioney Gomes, University of Brasília

    Professor Titular do Departamento de Linguística, Português e Línguas Clássicas da Universidade de Brasília (UnB). Pesquisa línguas indígenas, português do Brasil e língua brasileira de sinais (Libras). Atua também na formação inicial e continuada de professores. Concluiu mestrado e doutorado em Linguística na UnB, tendo sido, durante este último período de formação, pesquisador visitante nos seguintes centros de pesquisa franceses: Centre d'Études de Langues Indigènes d'Amérique (CELIA/Paris) e Laboratoire Dynamique du Langage (DDL/Lyon). Foi coordenador do Programa Institucional de Bolsas de Iniciação à Docência (PIBID/CAPES) do curso de Letras (2014-2018) e coordenou o Programa de Pós-graduação em Linguística da UnB (mestrado e doutorado) no biênio 2012-2013. É líder do grupo de pesquisa Funcionalismo, Tipologia e Ensino (CNPq) e membro-fundador do grupo de pesquisa Núcleo de Tipologia e Línguas Indígenas (NTL/CNPq). Juntamente com a Profa. Dra. Alejandra Regúnaga (CONICET e UNLPam, Argentina), coordena o Projeto 9 "Diversidade linguística na América (Línguas Ameríndias)" na Associação de Linguística e Filologia da América Latina (ALFAL). É membro-fundador da Rede de Investigação e Cooperação Interinstitucional sobre Diversidade Linguística (RICIDIL), a qual reúne universidades do Brasil, México, Argentina e Chile. Foi coordenador de Pesquisa e Inovação do Instituto de Letras da UnB (2022). Coordena o PIBID Letras Português desde junho de 2024.

References

ALENCAR, L. F. Yauti: A Tool for Morphosyntactic Analysis of Nheengatu within the Universal Dependencies Framework. Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL). Anais... In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL). SBC, 25 set. 2023. Disponível em: https://sol.sbc.org.br/index.php/stil/article/view/25445. Acesso em: 28 jun. 2024.

ALENCAR, L. F. A Universal Dependencies Treebank for Nheengatu. Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 2. Anais... In: PROPOR 2024. Universidade de Santiago de Compostela, Galicia, Spain: Association for Computational Linguistics (ACL), 2024. Disponível em: https://aclanthology.org/2024.propor-2.0.pdf. Acesso em: 28 jun. 2024.

BIRD, S.; KLEIN, E.; LOPER, E. Natural language processing with Python: Analyzing Text with the Natural Language Toolkit. 2019. Disponível em: https://www.nltk.org/book/. Acesso em: 15 jul. 2024.

BRUNATO, D.; CIMINO, A.; DELL’ORLETTA, F.; MONTEMAGNI, S.; VENTURI, G. Profiling–UD: a Tool for Linguistic Profiling of Texts. Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), 2020, p. 7145-7151. Disponível em: http://www.lrec-conf.org/proceedings/lrec2020/pdf/2020.lrec-1.883.pdf. Acesso em: 15 jul. 2024.

CIERI, C.; MAXWELL, M.; STRASSEL, S.; TRACEY, J. Selection Criteria for Low Resource Language Programs. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16). Portorož: European Language Resources Association (ELRA), 2016. p. 4543-4549. Disponível em: https://aclanthology.org/L16-1720/. Acesso em: 15 jul. 2024.

HÄMÄLÄINEN, M. Endangered Languages are not Low-Resourced. Preprints 2021. Disponível em: https://doi.org/10.20944/preprints202104.0113.v1. Acesso em: 15 jul. 2024.

JURAFSKY, D.; MARTIN, J. H. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. 3 ed. 2023. Disponível em: https://web.stanford.edu/~jurafsky/slp3/. Acesso em: 15 jul. 2024.

KARGARAN, A. H.; IMANI, A.; YVON, F.; SCHÜTZE; H. GlotLID: Language Identification for Low-Resource Languages. Findings of the Association for Computational Linguistics: EMNLP 2023, December 6-10, p. 6155-6218. Disponível em: https://aclanthology.org/2023.findings-emnlp.410.pdf. Acesso em: 15 jul. 2024.

LI, Xiuhong; LI, Zhe; SHENG, J.; SLAMU, W. Low-Resource Text Classification via Cross-lingual Language Model Fine-tuning. Proceedings of the 19th China National Conference on Computational Linguistics, 2020 p. 994-1005. Disponível em: https://aclanthology.org/2020.ccl-1.92.pdf. Acesso em: 17 jul. 2024.

LINDERS, G. M.; LOUWERSE, M. M. Lingualyzer: A computational linguistic tool for multilingual and multidimensional text analysis. Behav Res, 2023. Disponível em: https://doi.org/10.3758/s13428-023-02284-1. Acesso em: 17 jul. 2024.

MAGER, M.; GUTIERREZ-VASQUES, X.; SIERRA, G.; MEZA, I. Challenges of language technologies for the indigenous languages of the Americas. Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, New Mexico, USA, August 20-26, p. 55-69. 2018. Disponível em: https://aclanthology.org/C18-1006/. Acesso em: 10 jul. 2024.

MANNING, C.; SCHÜTZE, H. Foundations of statistical Natural language Processing. Cambridge: MIT Press, 1999.Disponível em: http://nlp.stanford.edu/fsnlp/. Acesso em: 16 jul. 2024.

RUETER, J.; FREITAS, M. F. P.; FACUNDES, S. S.; HÄMÄLÄINEN, M.; PARTANEN, N. Apurinã Universal Dependencies Treebank. Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas. Anais... In: PROCEEDINGS OF THE FIRST WORKSHOP ON NATURAL LANGUAGE PROCESSING FOR INDIGENOUS LANGUAGES OF THE AMERICAS. Online: Association for Computational Linguistics, 2021. Disponível em: https://www.aclweb.org/anthology/2021.americasnlp-1.4. Acesso em: 30 nov. 2023.

SANDALO, M. F. S.; GALVES, C. M. C. Anotando sintaticamente uma língua originária do Brasil: o problema de Anchieta. Cadernos de Estudos Linguísticos, v. 65, p. e023007–e023007, 2023.

TEIXEIRA DE SOUSA, L. Sobre a constituição de corpora para línguas com poucos recursos. Revista Linguíʃtica, v. 16, n. 1, p. 43-61, 30 abr. 2020.

THOMAS, G. Universal Dependencies for Mbyá Guaraní. (A. Rademaker, F. Tyers, Eds.). Proceedings of the Third Workshop on Universal Dependencies (UDW, SyntaxFest 2019). Anais... In: UDW-SYNTAXFEST 2019. Paris, France: Association for Computational Linguistics, ago. 2019. Disponível em: https://aclanthology.org/W19-8008. Acesso em: 10 jun. 2024.

VIRK, S. M.; FOSTER, D.; MUHAMMAD, A. S.; SALEEM, R. A Deep Learning System for Automatic Extraction of Typological Linguistic Information from Descriptive Grammars. Proceedings of Recent Advances in Natural Language Processing, 2021, p. 1480-1489. Disponível em: https://aclanthology.org/2021.ranlp-1.166.pdf. Acesso em: 10 jun. 2024.

Published

2026-08-18

How to Cite

Linguística Computacional e línguas indígenas brasileiras. PERcursos Linguísticos, v. 17, n. 40, 18 Aug.2026.