Lab in the loop: teaching an AI what toxic looks like
Most companies training an AI model start with the data they already have. Tom Masterson thinks that is usually the wrong place to start.
Masterson is chief operating officer of Deep Genomics, a Toronto company that uses AI to discover genetic medicines. On Between Two COOs he walked through how the company builds its training data, and why it deliberately tests molecules that are toxic.
What a lab in the loop is
A lab in the loop is an iterative cycle. You design an experiment, run it, analyze the result, adjust, and feed what you learned back into the model. The model then shapes the next round of experiments.
Masterson compares it to lean methodology in software: rapid iteration drives better performance. At Deep Genomics the lab is where all the experimentation happens that creates the data for the company's models.
His first point is that the loop itself is not the advantage. "Anybody can build a lab in the loop, which is just an iterative training cycle," he said. The advantage is in what goes into it.
Precision, not creativity
Most AI products are rewarded for being generative. Did it come up with something new? Was it creative? In drug development the reward is different.
"If you are not delivering the correct result, you kill the patient," Masterson said. A model that invents a clever new answer is not a feature there. Everything is precision oriented, and that changes what good training data looks like.
He said the company wants data it can be confident in: clear provenance, causal in nature, rigorously curated, and designed for machine learning from the start.
Why historical data trains badly
This is the counterintuitive part. When Deep Genomics looked at the historical data a pharma company might hand over, it was biased.
The reason is simple once you hear it. For decades, drug companies have tried to design molecules that are effective and safe. So their records are full of molecules that were already close to working. A model trained on that never sees enough toxic molecules to learn what toxic looks like.
So Deep Genomics designs test molecules that are intentionally toxic, "so that the model learns what toxic looks like," as Masterson put it. He said the data needs a certain amount of noise, probably more noise than good signal. Its data sets were more than 10 times larger than those of the partners it worked with, and designed in a much better way.
The result it produced
That approach fed the siRNA model behind the result that opens the episode. Given a cholesterol target it had never seen and about 3,500 candidate molecules, the model picked nine. One of them is a drug already approved and in patients.
Masterson is careful to call it a retrospective study. What he emphasizes is that the model had never seen that gene. "What that tells us is that the model has actually learned the rules, not that it's memorized the answers to some test."
What operators can take from this
You probably do not run a wet lab. The lesson still carries over to any team training or tuning a model on company data.
Your historical data reflects the decisions you already made. A sales model trained only on won deals, or a hiring model trained only on people you hired, has the same blind spot as a drug model trained only on safe molecules. It never learns what failure looks like.
Masterson's other rule is cultural. He tells his engineers to pretend every line of code is going into a patient, because some day it might. Most operators are not shipping drugs. The discipline still helps: know where every piece of training data came from, and whether it shows cause or only correlation.
FAQ
What is a lab in the loop? A lab in the loop is an iterative cycle where lab experiments generate data that trains an AI model, and the model's predictions shape the next experiments. Tom Masterson of Deep Genomics describes it as design, test, analyze, and adjust, repeated.
Why does Deep Genomics test molecules that are intentionally toxic? So the model learns what toxic looks like. Masterson says pharma's historical data is biased toward safe, effective molecules, so Deep Genomics deliberately designs toxic molecules to give its models the missing signal.
Is historical company data good for training AI models? Often not on its own. Masterson says historical data reflects past decisions and is biased as a result. Data designed for machine learning, with clear provenance and causal structure, trains better.
What did Deep Genomics' siRNA model find? Given a cholesterol target it had never seen and about 3,500 candidate molecules, the model picked nine, and one of them is a drug already approved and in patients. Masterson calls it a retrospective study.
What does lab in the loop mean for operators outside biotech? Any team training a model on its own records should ask what the data leaves out. A model that only sees your successes will not learn to recognize failure.
The COO's Execution Playbook
Frameworks, templates, and hard-won lessons from operators who've been in the chair. Every Tuesday.
No spam. Unsubscribe anytime.