Research

My research is motivated by inference problems in which observations are incomplete, noisy, structured, or generated by several competing mechanisms. I use tools from probability, statistics, optimisation, and machine learning to understand what information is present in the data, what assumptions make inference possible, and how these ideas can be translated into computational methods. This perspective began with population-genetic models and now informs my work on language, healthcare, scientific imaging, signal processing, and AI systems.

previous

Population genetics and demographic inference

Problem. Genomic data often look similar under very different demographic stories, which makes it hard to tell population-size change apart from population structure.

Approach. I use coalescent theory, stochastic processes, identifiability analysis, and simulation-based inference, including the inverse instantaneous coalescence rate (IICR).

Domains. Population genetics, evolutionary biology, genomic data analysis.

Associated publications are listed on the publications page; project pages link the relevant records.

current

NLP, embeddings, and topic modelling

Problem. Large text collections are difficult to summarise in a way that remains interpretable and faithful to the underlying corpus.

Approach. I compare classical topic models such as LDA and NMF with embedding-based methods, including transformer encoders, coherence measures, and divergence statistics.

Domains. Scientific and technical text, safety-report analysis at a non-confidential level, and student research in representation learning.

Associated publications are listed on the publications page; project pages link the relevant records.

current

AI for healthcare and scientific imaging

Problem. Clinical and scientific images are high-dimensional, sensitive, and often incomplete, which raises questions of representation, uncertainty, and privacy-aware use.

Approach. I work on segmentation, generative models, decision-support workflows, and deployment constraints rather than on product claims or partner-specific architectures.

Domains. Medical imaging and scientific data analysis.

Associated publications are listed on the publications page; project pages link the relevant records.

current

Hybrid models, signal processing, and agentic systems

Problem. Learned components are often most useful when combined with classical models, filters, or tools rather than used in isolation.

Approach. I study combinations of neural models with model-based filtering, and, more recently, LLM agents, tool use, retrieval, and orchestration.

Domains. GNSS positioning as one example, together with domain-specific language systems. Agentic AI is a current technical direction, not the historical centre of this research profile.

Associated publications are listed on the publications page; project pages link the relevant records.

Research timeline

  1. 2012–2016

    PhD in Applied Mathematics and Statistics

    INSA Toulouse / Institut de Mathématiques de Toulouse

    Probabilistic modelling of demographic history from genomic data, with emphasis on coalescent processes, population structure, and identifiability.

  2. 2016–2017

    Research engineer

    INRAE

    Genetic linkage maps and related computational genetics.

  3. 2017–2019

    Research engineer

    Institut de Mathématiques de Toulouse

    Probabilistic modelling in genetics.

  4. 2019–2021

    Research engineer

    ENAC

    Natural-language processing for aviation safety, described here only at a non-confidential level.

  5. 2023–present

    AI Researcher

    Torus

    Medical imaging, signal processing, generative models, and agentic systems, presented as selected applied research rather than commercial case studies.

See also selected projects and the 2016 PhD defence archive.