AI presents a unique promise to dramatically improve human health by driving transformative progress in medicine. The world is working to advance the fundamental technology of AI, but between the abilities of large language models and the daily work of discovering drugs is a chasm. Mirror is building the bridge.
The pharmaceutical industry represents the technological engine through which we most directly extend and improve human lives. Medicine eradicated smallpox. It turned HIV from a death sentence into a daily pill, made childhood leukemia survivable in nine out of ten cases, and averted millions of deaths in the last global pandemic. Medicine creates hope where there’s otherwise little to be found.
But developing drugs is hard. The rise in time and expense has been extensively documented1: For nearly half a century, we’ve collectively stagnated at one of humanity’s most important undertakings, despite unimaginable advancements in technology.
Artificial intelligence creates new possibilities for advancing medicine. For a field fundamentally entangled with the complexity of human biology, the emergence of expert intelligence deployable at unprecedented speed and scale seems an obvious boon. Yet the capital hasn’t followed: in 2025, AI drug discovery garnered roughly ten billion in venture funding, while just four hyperscalers committed hundreds of billions to datacenters alone.2
As AI consumes an ever growing share of the global economy, it bears an ever larger burden to create immense value for society. A technology absorbing energy and capital at this scale should be aimed at problems of commensurate importance. Solving medicine is the singular problem critical enough to justify that investment, let alone deliver on it.
At Mirror, we see this as a mandate. Improving health through AI isn’t just an economic opportunity; it’s an obligation, one we share a responsibility to address as quickly as possible.
Efforts to wield large language models to drive discovery are filtering through every discipline of science.
Already, proofs of concept have shown how LLMs can manipulate, contemplate and connect knowledge across vast and disparate domains. They’ve been leveraged to plan new reaction pathways for synthesizing chemicals; to process qualitative measurements buried in technical papers into structured datasets for materials science; to propose new hypotheses in biology and analyze experimental data.1, 2, 3
As the low-hanging fruit is picked off by well-deployed models, those building AI for science will have to face the limitation that in general, new scientific truths are not contained entirely within model weights. LLMs might have the ability to represent arbitrary conclusions (assuming they can be represented by human language), but that doesn’t mean the model alone is sufficient to arrive at those truths in practice, let alone substantiate them.
Fortunately, LLMs don’t need to be an omniscient oracle to be useful. Humanity already developed a wildly successful engine for discovery: the scientific method. Instead of asking models to answer new scientific questions directly, we should deploy them to accelerate the practice of science; that is, to facilitate scientists’ efforts in investigating those questions. This is a core part of our thesis on agentic drug discovery: The fastest way to create AI that delivers breakthroughs in the clinic is to build directly for the humans that are already translating scientific discoveries into novel medicines.
That’s why Mirror’s AI platform is singularly focused on supercharging pharma scientists. By grounding our efforts in the challenges faced at the real-world frontier, we ensure each capability increase translates to tangible insight and acceleration for practitioners pushing programs forward. The measure of progress switches from what the model can do, to what scientists can do with it.
For agents to be effective accelerants, raw intelligence isn’t enough.
The practice of pharmaceutical companies involves deep specialization across drug modalities and stages of the development pipeline. Day-to-day decisions can depend entirely on a program’s history, an organization’s practices, on a team’s composition and constraints. And every company has unique technology and data that is critical to its edge.
Just like employees need onboarding, a practical agent needs to understand and inhabit the ecosystem in which it will work. That starts with acquiring context: program objectives, target/product hypotheses, prior art, etc. But context alone isn’t enough; agents need an execution environment with regimented access to an org’s software, data, and workflows. And, of course, they’ll need to have highly specialized skills and expertise in drug discovery tasks.
Strong AGI believers might misunderstand this specialization as futile disregard for The Bitter Lesson.1 But general intelligence doesn’t eliminate the need for systems that humans can understand, control, and embed within real organizations. The engineering to translate frontier capabilities into solutions that actually fit customer needs isn’t a detour; it’s critical infrastructure.2
The scientists driving preclinical programs are extraordinarily busy. They’re designing biochemical assays, purifying synthesized compounds, programming liquid handlers, designing knockout screens, analyzing dose-response data, reviewing selectivity profiles, drafting regulatory documents — the list could go on. They need agents that come ready to work. Like a good employee, agents should save scientists time on day one, and get better as scientists invest time into them.
Building the AI systems to meet this need represents an exciting, valuable and long-tailed engineering challenge.
Agents improve at tasks on which we can reliably verify success. They iteratively attempt tasks, assess which strategies succeed, and then optimize against the successes. The hard part is knowing what to improve on, and how to verify success — i.e., developing a benchmark.
If agents are going to deliver tangible benefits, they’ll need to be heavily focused on addressing user needs. That means deeply understanding the demands of drug discovery scientists, and working alongside them to ensure benchmark tasks and verifiers encapsulate their needs.
Tasks can be hierarchical and interdependent; some are granular and some are open-ended. Plus, the work of drug discovery is long-horizon and path-dependent. Early mistakes have downstream consequences, and the ultimate source of verification is always experimental. A good benchmark will effectively catalogue, distinguish and prioritize these tasks.
When designed well, benchmarks can compound progress. The better agents get, the harder the problems they can tackle. Driving improvements in speed, consistency, and cost-efficiency on a set of essential tasks furthers the frontier of work agents can perform, exposing new capability gaps and new opportunities to improve.
User-driven benchmarks are a central focus of Mirror’s development strategy. They serve as the roadmap for our AI platform, and present a scalable path towards ever-improving agents.
Scientists have far more ideas to explore than resources or bandwidth to do so. A single pharma program could consider hundreds of possible target-disease pairs, numerous modalities and innumerable compounds, various optimization strategies, etc., and somewhere in that jungle of possibilities lie relatively few life-saving combinations. Discovery is bottlenecked by the time and capital limits of exploration.
Despite the infinite possibilities, expert scientists have a remarkable ability to weigh chemical and biologic evidence and intuition to guide decisions across each stage of the development pipeline and discover the medicines we rely on. The fact that this discovery is accomplished in teams demonstrates that experts are willing to delegate consequential work when they can understand and trust how it’s being done.
That’s a critical bar for AI to meet. For an agent, trust translates to consistency and transparency. What did it do, how, and why? It did something smart the last nine times I checked—how about the tenth? If an agent can demonstrate reliable performance while transparently grounding its decisions in auditable data and tool-use, then there’s a path to establishing trust.
The importance of efficiency is more straightforward. If we’re to search more paths through the space of potential therapeutics, we’ll need to minimize the cost per search.1 Cost optimizations can come from across the AI stack, through more efficient inference and software, better context curation and model selection, and economies of scale.
Armed with efficient, trustworthy agents, human experts can explore a far wider range of hypotheses in parallel, maintaining visibility and control without the overhead of executing each exploration one-at-a-time. In that world, the quality of clinical candidates should rise significantly — with many hypotheses competing to advance, the bar for entrants moves from minimally viable to truly exceptional.
Enormous cost and uncertainty stand between promising biological ideas and treatments that reach patients. Making drug discovery dramatically faster and cheaper would usher in a wave of new therapeutics, and render entire classes of therapeutics economically viable, including treatments for rare diseases and personalized medicine.
Creating that future will require more than advances in preclinical development—we’ll need breakthroughs in clinical performance. The medicines that matter are the ones that prove safe and effective across patients in clinical trials, and predicting the clinical performance of candidates remains a major challenge, let alone directly designing strong performers.
Hard inverse problems like this are often solved by first making the forward problem tractable. Designing a therapy directly for clinical efficacy is extraordinarily difficult, in large part because the pipeline from molecular intervention to human outcome remains slow, expensive and unreliable. As AI accelerates each component of that pipeline, clinical candidates become cheap and abundant, making the design problem increasingly approachable. Orders-of-magnitude improvements in the preclinical pipeline could create the data, iteration speed, and economic freedom required to eventually optimize medicines for clinical performance.
We see the agent systems we’re building today as a launchpad toward that future. The immediate opportunity is to give scientists radically more leverage to accelerate preclinical drug discovery. The stretch goal is to compress the entire path—from selecting a disease to obtaining evidence of an intervention that improves human health—into a matter of months.
There’s an exciting road ahead; Mirror’s just getting started. We’re looking for unique and driven individuals who share our vision. If that’s you, reach out to us.