Research
Three threads, one gap: between what a sequencer reports and
what a biologist can defend.
01
Making transcriptome quality measurable
Every long-read pipeline reports a different number of isoforms, and none of them can tell you which ones are wrong.
The problem
A real isoform and an artefact look identical in a long-read catalogue. Without a ground truth, "40,000 novel isoforms" is a claim, not a result.
What I built
TUSCO — genes with a single known isoform, so any extra call at one of them is a false discovery by construction. Quality becomes a number. It is now the evaluation layer inside SQANTI3.
02
Learning chromatin state from single molecules
Structure and regulation are usually measured on different molecules and correlated afterwards.
The problem
A catalogue says which isoforms exist, not why one is expressed. Native long reads carry methylation, nucleosomes and occupancy on the same molecule as the transcript.
What I built
SQANTI-epi — one coordinate system for all of it, with a machine-learning model scoring accessibility per molecule. It is judged under a protocol frozen before the numbers are seen.
03
Making AI interpretation accountable
A language model will write you a beautiful biological story. That is exactly the problem.
The problem
The last step of an omics study is the least controlled. A gene list becomes a narrative, judged on whether it sounds plausible rather than on the evidence under it.
What I built
TELLME and PaintOmics AI — agents that cite what they claim, keep only the links the data confirms, and delete every unsupported gene, PMID and functional claim before you see it.