teodora vuković

Making the whole human visible.

I build AI platforms and research infrastructure that integrate voice, gesture, language, and physiology, advancing the science of human signal at scale. I work in the open, so that what we learn about human beings remains a shared resource.

Multimodal AI Human Signal Processing Open Science
Flagship Project ↓
Teodora Vuković Portrait
Flagship Project

MOSAIC

Multimodal Open-Science AI-Powered Integrative Computation

Humans don't communicate in channels. They speak, move, breathe, and feel simultaneously. MOSAIC is an AI-powered data science platform for integrating physiological, behavioral, linguistic, clinical, and subjective data into a unified research infrastructure. Built for science, designed for the full complexity of the human signal. It is open by design: its methods, standards, and tooling are released to the wider research community, because infrastructure for understanding people should serve the public, not be held privately.

Innosuisse Innovation Booster AI Open Science Human Signal Processing University of Zurich
See All Research ↓

Engagement

Current

2023 – Present

Multimodal Technology Group Leader

University of Zurich (UZH), Department of Computational Linguistics

Guiding research, supervising junior researchers and teaching in the domain of multimodal language and behavior processing using video, speech and natural language processing (NLP).

University of Zurich
2026 – Present

MOSAIC – Project Coordinator & Principal Investigator

Multimodal Open-Science AI-Powered Integrative Computation

Innosuisse Innovation Booster Artificial Intelligence

AI-powered data science platform for processing human signal, involving physiological, behavioral, linguistic, clinical, and subjective data.

2026

LaDaD – Principal Investigator & Lecturer

Language Data Across Disciplines

CLARIN-CH / Linguistic Research Infrastructure (LiRI), University of Zurich

Development of teaching materials and workshops on data analytics and quantitative and computational skills for language data, run jointly across nine Swiss institutions. The full programme is documented on the CLARIN-CH project page.

2025 – 2027

DAVA Lab – Project Lead

Digital Audio-Visual Annotation Lab

Digital Society Initiative (DSI) / Linguistic Research Infrastructure (LiRI), University of Zurich

Development of AI-powered software for automatic audio-visual and multimodal data analysis, aimed at academia and industry. The lab and its aims are described on the Digital Society Initiative pages.

2025 – 2027

MAVA Cluster – Principal Investigator

Multimodal Audio-Visual Analysis Cluster

Swiss Data Science Center (SDSC), ETH Zürich / Linguistic Research Infrastructure (LiRI), UZH / Film Studies (FiWi), UZH / TIB Hannover

Cluster for infrastructure interoperability and data exchange, connecting VideoScope, TIB-AV-A and VIAN through shared formats and APIs so that audio-visual research data can move between them under FAIR principles. The resulting exchange format for video annotations is developed in the open as mava-exchange, and the cluster is described in the announcement at LiRI.

2023 – 2027

CAPIRE – Project Coordinator & Principal Investigator

Computational Analysis of Personal Identity in Interaction: Recognition and Ethics

Digital Society Initiative (DSI) / Department of Computational Linguistics, UZH

Focus on explainable machine learning for personal identity patterns in speech, facial expressions and gestures.

Past

2021 – 2025

LiRI Corpus Platform Project Management

Linguistic Research Infrastructure (LiRI), University of Zurich

Application and research methodology development, implementing and training models for image and speech recognition, natural language processing (NLP). The platform is live at lcp.linguistik.uzh.ch — with CatchPhrase, SoundScript and VideoScope for text, audio and video corpora — and developed in the open on GitHub.

2017 – 2021

TraCeBa Project – Electronic Language Resources

(Dis-)entangling Traditions on the Central Balkans: Performance and Perception

Slavisches Seminar (Department of Slavic Studies), University of Zurich

Creation of electronic language resources and corpus-based analysis of linguistic variation. The resulting Spoken Torlak Dialect Corpus — semi-orthographic transcripts of 86.5 hours of fieldwork recordings — is openly published through CLARIN.SI, where it can also be queried directly in the browser. Funded by ERA.Net RUS Plus under FP7/Horizon 2020.

University of Zurich Slavisches Seminar European Union ERA.Net RUS Plus
2016 – 2017

PhD Researcher

Text Group, Language and Space Lab, University of Zurich (UZH)

Developing the corpus of the Torlak dialect, collected in the field across the Timok area and later released as the openly available Spoken Torlak Dialect Corpus, searchable via CLARIN.SI.

University of Zurich Language and Space Lab
2012 – 2016

Research Associate

Institute for Balkan Studies, Serbian Academy of Sciences and Arts (SASA)

Corpus linguistics, digital humanities, fieldwork, dialectology, digital lexicography. Author of the Bunjevac dialect corpus.

Meditation & Contemplative Research

Practice and teaching alongside research that turns the same multimodal methods — language, speech, physiology — toward first-person experience.

2020 – Present

meditative.dev – Co-founder & Meditation Teacher

Meditation practice and research initiative

A human development initiative grounded in truth seeking and the reduction of suffering, teaching a variety of meditation tools in the spirit of pragmatic dharma. More at meditative.dev.

2026 – Present

Meditative Phenomenology in Fire Kasina Retreats

With Prof. Dr. Marjorie Woollacott and Dr. Niffe Hermansson

NLP analysis of personal experience reports and meditative phenomenology reports gathered during an advanced meditation retreat.

2024, 2025

Fire Kasina Meditation and Oscillatory Network Dynamics

Riddle Lab, Florida State University — Principal Investigator: Prof. Justin Riddle

Participant in an advanced meditation retreat during which the Riddle Lab collected EEG and speech data.

2024

Real-time Biofeedback Meditation Robot

Waseda University, Tokyo — Principal Investigator: Prof. Gabriele Trovato

Meditation guidance and multimodal analysis expertise for a robotic system delivering personalized meditation instruction through neurofeedback.

2022 – 2024

The Big Data Project

Emergent Phenomenology Research Consortium — Principal Investigator: Prof. Shiri Dori-Hacohen, University of Connecticut

NLP, AI and statistical analysis for an exploratory project investigating the process of awakening through large text collections and social network data. Member of the Emergent Phenomenology Research Consortium, a multidisciplinary alliance studying emergent phenomena with scientific and clinical methods.

Publications

170 citations · h-index 7 · i10-index 5 (Google Scholar, August 2026)
Full profile on Google Scholar →

In review

  • A Modular Approach Towards Representing Multi-Dimensional Corpora: Theoretical Considerations behind the LiRI Corpus Platform / Linguistic Data Environment

    Teodora Vuković, Jonathan Schaber, Johannes Graën, Jeremy Zehr, Igor Mustač, Daniel McDonald, Nikolina Rajović, Gerold Schneider, Noah Bubenhofer

    Submitted to Language Resources and Evaluation

2026

  • Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing

    Masoumeh Chapariniya, Jean-Marc Odobez, Volker Dellwo, Teodora Vuković

    Proceedings of FG2026, IEEE Biometrics

  • Spoken Corpora

    Teodora Vuković

    International Encyclopedia of Language and Linguistics, 3rd edition. Elsevier

  • Towards Language-Independent Face-Voice Association with Multimodal Foundation Models

    Aref Farhadipour, Teodora Vuković, Volker Dellwo

    Odyssey 2026

  • TidyVoice 2026 Challenge Evaluation Plan

    Aref Farhadipour, Jan Marquenie, Srikanth Madikeri, Teodora Vuković, Volker Dellwo, Kathy Reid, Francis M. Tyers, Ingo Siegert, Eleanor Chodroff

    Interspeech 2026

2025

  • The Swiss FAIR-compliant ecosystem of infrastructures 2.0

    Cristina Grisot, Alexandru Craevschi, Christian Futter, Teodora Vuković, Jeremy Zehr, Julia Krasselt, Philipp Dreesen

    CLARIN Annual Conference Proceedings

  • Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics 2nd Best Paper

    Masoumeh Chapariniya, Teodora Vuković, Sarah Ebling, Volker Dellwo

    15th International Conference on Computer and Knowledge Engineering (ICCKE 2025), IEEE

  • LiRI Corpus Platform: Demonstration of a Web-Based Infrastructure for Multimodal Corpus Analysis

    Teodora Vuković, Jeremy Zehr, Jonathan Schaber, Johannes Graën, Igor Mustač, Daniel McDonald, Nikolina Rajović, Noah Bubenhofer

    Interspeech 2025

  • NLP for preserving Torlak, a vulnerable low-resource Slavic language

    Li Tang, Teodora Vuković

    Proceedings of the 31st International Conference on Computational Linguistics. ACL Anthology

  • CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation

    Aref Farhadipour, Shiran Liu, Masoumeh Chapariniya, Valeriia Vyshnevetska, Srikanth Madikeri, Teodora Vuković, Volker Dellwo

    NIST SRE 2024

  • Team Switzerland Submission to NIST SRE24 Speaker Recognition Evaluation

    Amrutha Prasad, Hatef Otroshi Shahreza, Andrés Carofilis, Aref Farhadipour, Shiran Liu, Srikanth Madikeri, Anjith George, Petr Motlicek, Sébastien Marcel, Masoumeh Chapariniya, Valeriia Perepelytsia, Teodora Vuković, Volker Dellwo

    NIST SRE 2024

2024

  • Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification

    Aref Farhadipour, Masoumeh Chapariniya, Teodora Vuković, Volker Dellwo

    ICNLSP 2024. ACL Anthology

  • Empirical approaches to variation: the case of the Timok variety of Torlak

    Teodora Vuković

    Doctoral dissertation, University of Zurich

  • The LiRI Corpus Platform

    Johannes Graën, Jonathan Schaber, Daniel McDonald, Igor Mustač, Nikolina Rajović, Gerold Schneider, Teodora Vuković, Jeremy Zehr, Noah Bubenhofer

    CLARIN2023: Selected papers. Leuven: CLARIN

2023

  • A Corpus-Based Analysis of the Grammatical Status of Short Demonstratives in the Timok Dialect

    Teodora Vuković

    Journal of Slavic Linguistics 31(1–2), 245–269

  • Toward Sociolinguistic Corpora of Torlak

    Maja Miličević Petrović, Teodora Vuković, Mirjana Mirić, Daria V. Konior, Anastasia Escher

    Zeitschrift für Slavische Philologie 79(1), 123–151

  • Under the Magnifying Glass: Dimensions of Variation in the Contemporary Timok Variety

    Teodora Vuković, Mirjana Mirić, Anastasia Escher, Svetlana Ćirković, Maja Miličević Petrović, Andrey N. Sobolev, Barbara Sonnenhauser

    Zeitschrift für Slavische Philologie 79(1), 153–194

2022

  • Degrees of non-standardness: Feature-based analysis of variation in a Torlak dialect corpus

    Teodora Vuković, Anastasia Escher, Barbara Sonnenhauser

    International Journal of Corpus Linguistics 27(2), 220–247

  • MultextEast V6 Torlak Specifications

    Teodora Vuković

    In Tomaž Erjavec (ed.), MultextEast V6. CLARIN.SI

2021

2020

  • Corpus-based analysis of spoken narratives: introducing a corpus and a search tool

    Philipp Wasserscheidt, Marija Mandić, Nadine Vollstädt, Ana Jovanović, Ivana Tanasijević, Teodora Vuković, Ivana Vučina Simović, Uliana Yazhinova, Anđelka Zečević

    Govor 37(2), 149–178. Zagreb: Hrvatsko filološko društvo

  • Automatic language profiling of a dialect speaker: the case of the Timok variety spoken in Berčinovac, Eastern Serbia

    Anastasia Makarova, Daria V. Konior, Teodora Vuković, Andrej N. Sobolev, Olivier Winistörfer

    Acta Linguistica Petropolitana 16(2), 160–180 (in Russian)

2019

  • Corpora and Processing Tools for Non-Standard Contemporary and Diachronic Balkan Slavic

    Teodora Vuković, Nora Muheim, Olivier Winistörfer, Ivan Simko, Anastasia Makarova, Sanja Bradjan

    Proceedings of the Student Research Workshop (RANLPStud 2019), RANLP 2019, Varna, 62–68

2018

  • Prostorna raspodela frekvencije post-pozitivnog člana u timočkom govoru [Areal distribution of the frequency of the post-posed article in the Timok vernacular]

    Teodora Vuković, Tanja Samardžić

    In S. Ćirković (ed.), Timok: folkloristička i lingvistička terenska istraživanja 2015–2017, 181–199

2017

  • Creation and some ideas for classroom use of an electronic corpus of the dialect of Bunjevci

    Teodora Vuković, Maja Miličević

    In J. Filipović & J. Vučo (eds.), Minority Languages in Education and Language Learning. Belgrade: Faculty of Philology

  • Digitalna zaštita bunjevačkih govora [Digital protection of the Bunjevac dialect]

    Biljana Sikimić, Teodora Vuković

    In D. Njegovan (ed.), Kultura i identitet Bunjevaca. Novi Sad: Muzej Vojvodine, 473–489

Work With Me

Interested in collaboration or consulting? Let's connect.