Marcos Zampieri

Assistant Professor
School of Computing
George Mason University
Fairfax, VA, USA
Headshot

About

I am an Assistant Professor at the School of Computing at George Mason University and currently a Visiting Assistant Professor at the Department of Computer Science at Duke University.

My research interests are in Computational Linguistics and Natural Language Processing (NLP), a core area of Artificial Intelligence (AI). I take a linguistically oriented approach to NLP with the goal of enhancing our understanding of human language and communication while improving the robustness, safety, and accessibility of NLP systems.

Here are some questions driving my research:

  • Language variation: How do we design corpora, benchmarks, and models that capture individual (e.g., speakers' L1 and proficiency) [NAACL 2024] [EMNLP 2025] and systemic (e.g., diatopic and diachronic) [COLING 2024] [NLP4DH 2026] variation, and what linguistic insights do these models reveal [Frontiers 2023]?
  • Multilingual NLP: How can we build and evaluate systems that are robust across languages [NAACL 2025a], including low-resource languages [NAACL 2025b] [ACL 2025] and non-standard [EMNLP 2026] language input?
  • LLM safety and evaluation: As LLM-based systems become widespread, how do we evaluate their safety and readiness for real-world deployment in domains ranging from robotics [ICRA 2026] to healthcare [EACL 2026]?

I am particularly interested in how AI, especially LLMs, is changing education [SIGCSE 2025] and how we can design and deploy safe and pedagogically aligned NLP systems to help students learn language and computing.

I co-founded and have served as co-organizer of the VarDial workshop since 2014. I have served as Program Chair for SemEval (2025–2026), Tutorial Chair for ACL 2022, and Faculty Advisor for NAACL SRW 2024, alongside regular area chair and senior area chair roles at major NLP venues.


Recent Selected Publications

For a full list of publications please check Google Scholar.

SUNDER: Selective Unmasking for Text Understanding
Alphaeus Dmonte, Tharindu Ranasinghe, Marcos Zampieri
EMNLP (2026)

Large-scale Multilingual News Image Captioning with LLMs
Yuji Chen, Purushoth Velayuthan, Alistair Plum, Hansi Hettiarachchi, Saroj Basnet, Menan Velayuthan, Marcos Zampieri, Tharindu Ranasinghe
EMNLP (2026)

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments
Amirreza Payandeh, Anuj Pokhrel, Daeun Song, Marcos Zampieri, Xuesu Xiao
ICRA (2026) pdf

TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
ACL (2025) pdf

Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations
Poorvi Acharya, J. Elizabeth Liebl, Dhiman Goswami, Kai North, Marcos Zampieri, Antonios Anastasopoulos
EMNLP (2025) pdf

mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation
Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri
NAACL (2025) pdf

Bayelemabaga: Creating Resources for Bambara NLP
Allahsera Auguste Tapo, Kevin Assogba, Christopher M Homan, M. Mustafa Rafique, Marcos Zampieri
NAACL (2025) pdf

Large Language Models in Computer Science Education: A Systematic Literature Review
Nishat Raihan, Mohammed Latif Siddiq, Joanna CS Santos, Marcos Zampieri
SIGCSE (2025) pdf

Annotator Reliability Through In-Context Learning
Sujan Dutta, Deepak Pandita, Tharindu Weerasooriya, Marcos Zampieri, Christopher Homan, Ashiqur KhudaBukhsh
AAAI (2025) pdf

A Survey of Multimodal Sarcasm Detection
Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanojia, Yu Kong, Marcos Zampieri
IJCAI (2024) pdf

Language Variety Identification with True Labels
Marcos Zampieri, Kai North, Tommi Jauhiainen, Mariano Felice, Neha Kumari, Nishant Nair, Yash Bangera
LREC-COLING (2024) pdf

Native Language Identification in Texts: A Survey
Dhiman Goswami, Sharanya Thilagan, Kai North, Shervin Malmasi, Marcos Zampieri
NAACL (2024) pdf

Features of Lexical Complexity: Insights from L1 and L2 Speakers
Kai North, Marcos Zampieri
Frontiers in Artificial Intelligence (2023) url

Lexical Complexity Prediction: An Overview
Kai North, Matthew Shardlow, Marcos Zampieri
ACM Computing Surveys (2023) url


Books

Automatic Language Identification in Texts

Automatic Language Identification in Texts

Tommi Jauhiainen, Marcos Zampieri, Timothy Baldwin, Krister Lindén
Synthetisis Lectures on Human Language Technologies
Springer (2024)


Similar Languages, Varieties, and Dialects

Similar Languages, Varieties, and Dialects: A Computational Perspective

Marcos Zampieri, Preslav Nakov (Editors)
Studies in Natural Language Processing
Cambridge University Press (2021)


Last Updated: September 2026 | Template: Plain Academic