Skip to content

About me

I am a Senior Researcher at Qualcomm AI Research in Amsterdam.

I completed my PhD at Vrije Universiteit Amsterdam, advised by Jakub Tomczak and Max Welling. Previously, I worked with Microsoft AI4Science, Qualcomm AI Research, and NVIDIA.

Efficient LLM architectures

My current work explores latent reasoning, compressed KV caches, and sub-quadratic attention, with an emphasis on reducing the memory and computational cost of language models.

  • NeurIPS 2026

    The Key to Going Linear: Analysis-Driven Transformer Linearization

    Anna Kuzina, Paul N. Whatmough, Babak Ehteshami Bejnordi

    Paper ↗

  • ICLR 2026

    KaVa: Latent Reasoning via Compressed KV-Cache Distillation

    Anna Kuzina, Maciej Pióro, Paul N. Whatmough, Babak Ehteshami Bejnordi

    Paper ↗

  • TMLR 2024

    Hierarchical VAE with a Diffusion-based VampPrior

    Anna Kuzina, Jakub M. Tomczak

  • TMLR 2024

    Variational Stochastic Gradient Descent for Deep Neural Networks

    Haotian Chen*, Anna Kuzina*, Babak Esmaeili, Jakub M. Tomczak

    Paper ↗Code ↗

  • NeurIPS 2022

    Alleviating Adversarial Attacks on Variational Autoencoders with MCMC

    Anna Kuzina, Max Welling, Jakub M. Tomczak

    Paper ↗Code ↗

Timeline

2025–present

Research

Qualcomm AI Research

Senior Researcher · Amsterdam

Efficient LLM architectures

2020–2025

Education

PhD · Vrije Universiteit Amsterdam

Latent Variable Generative Models

Advised by Jakub Tomczak and Max Welling

2024

Research

Microsoft AI4Science

Associate Researcher · Amsterdam

Machine learning force fields

2023

Research

Microsoft AI4Science

Research Intern · Cambridge

Machine learning force fields

2021

Research

Qualcomm AI Research

Research Intern · Amsterdam

Generative priors for inverse problems

2017–2019

Education

MSc · HSE and Skoltech

Computer Science · Statistical Learning Theory

2018

Research

NVIDIA

Deep Learning Intern · Moscow

Object detection and tracking