Funding the Future of Science: Q&A with Dorothy Chou
As we push deeper into the frontiers of biology, climate, and materials, the challenges we face are increasingly complex and costly — yet the urgency to solve them has never been greater. Artificial intelligence offers a rare and powerful catalyst, but seizing that opportunity requires more than good models and clever algorithms. It requires the right funding, the right infrastructure, and the right partnerships.
At Renaissance Philanthropy, we build time-bound, thesis-driven philanthropic funds that advance fields. We believe that AI, large-scale datasets, and lab automation are converging in ways that could fundamentally accelerate the pace of scientific discovery. But that acceleration depends on open infrastructure: the shared tools, datasets, and foundational research that no single lab or company has the incentive to build alone. Supporting that infrastructure is one of our priorities.
Google.org and Google DeepMind are taking a significant step in that direction, having recently launched a new AI for Science funding strategy alongside a $30M global Impact Challenge, in partnership with Renaissance Philanthropy. The strategy sets out a bold vision for building a sustainable, high-impact complement to government funding for AI accelerated science.
To explore what this means for the future of science and discovery, Tom Kalil sat down with Dorothy Chou from Google DeepMind.
Tom Kalil: Let's start with the "why now”. AI for science isn't a new idea. What has changed that makes this the right moment to launch a dedicated funding strategy, and what gap does it address that existing funding mechanisms aren't filling?
Dorothy Chou: Two things have fundamentally changed. First, AI’s capability threshold in science shifted much faster than most people expected. Just a few years ago, applying frontier AI to scientific problems still required significant custom engineering for each specific domain. Now, we have models that can transfer knowledge meaningfully across biological, chemical, and physical systems, which changes the economics of discovery entirely. And AI does not appear to be slowing down, with multiagent AI for Science systems like Co-scientist suggesting that a new era of virtual scientific collaboration might be just around the corner.
This leads to a second shift: urgency of a different kind. Climate change, antimicrobial resistance, food security, these have always been crises. What’s changed is the cost of delay. We now have tools that can move at pace; every month lost to legacy funding structures is no longer just slow, it’s a systemic failure we can actually name.
The gap our strategy addresses is structural. Existing funding mechanisms were built for a different model of science: patient, linear, discipline-specific. What AI for science actually needs is interdisciplinary infrastructure, open data foundations, and validation pipelines that bridge the awkward gap between research grants and commercial investment. Nobody owns that space. This strategy is an attempt to name it clearly and attract funding toward it deliberately.
TK: AlphaFold is the obvious proof point here, but there's a risk people treat it as a once-in-a-generation stroke of luck, given that researchers started contributing to the Protein Data Bank in 1971, as opposed to a replicable model. What actually made AlphaFold work, and what would need to be true to replicate that success in other domains?
DC: AlphaFold worked because three things were true simultaneously. All three are replicable, but none of them are accidental.
First, a well-defined problem with a clear success metric. The CASP competition meant an expert community could agree on whether you were making progress. That sounds obvious, but many scientific problems don't have an equivalent.
Second, data availability at scale for training and validation. The PDB illustrates why this matters. Although it was built by scientists for sharing rather than machine learning, its decades of accessible records became the foundation for AlphaFold. That’s both a vindication of the instinct to share and a reminder of how much value is gained from open, collaborative data infrastructure.
Third, a research culture genuinely comfortable with long time horizons and indifferent to disciplinary boundaries. That's rarer than it sounds.
The replication question is whether those three conditions can be created in other domains, rather than stumbled into. I think they can. The fields where I'm most optimistic are those where the data exists but hasn't been aggregated, and where the scientific community is willing to agree on what good looks like. Genomics, climate modeling, materials discovery, parts of neuroscience. The strategy is designed to accelerate exactly that readiness work.
TK: The strategy introduces the idea of a "virtuous cycle", where philanthropic investment in open foundations lowers the barriers for startups, VC-backed success generates signals about where the gaps are, and those signals feed back into smarter philanthropic decisions. That's a compelling idea, but coordination between funders with very different incentives and timescales doesn't happen automatically. What would actually start that flywheel?
DC: Coordination problems don't solve themselves, and the virtuous cycle framing can become an excuse for passivity if you're not careful. "The ecosystem will figure it out" is usually what people say when they don't want to take responsibility for owning it.
What actually starts the flywheel is a small number of people who sit at the intersection of the different funding worlds and are willing to connect the dots between funding ecosystems that don't normally talk. That means genuinely learning each other's constraints and incentives, being explicit about what philanthropic capital has funded and what it's learned, and creating forums where VC signals can actually inform grant decisions rather than just running in parallel.
Google.org's Impact Challenge is partly designed to do exactly that. It's not just about the funding. It's about generating a visible, curated signal about where the genuinely AI-shaped problems are, what the failure modes look like, and where the gaps between prediction and validated result keep appearing. If we publish that learning honestly, it becomes useful to every funder in the space, not just Google.org.
TK: One of the hardest strategic questions in this space is deciding who funds what. Philanthropic capital, venture capital, and government funding each play a distinct role — but in practice the lines get blurry. How do you think about which parts of the AI for science stack are best suited to each type of funder, and where do hybrid approaches become necessary?
DC: The cleanest way I think about it is by time horizon and risk appetite.
Government funding is best suited to the long-horizon, pre-competitive layer: foundational datasets, shared infrastructure, scientific workforce development. These are things that benefit everyone and that no single investor can fully capture returns from. Public funding has largely abdicated this role in recent years, and the field is paying for that. Historically, the state has been a lead risk-taker for public good. When government (eg DARPA in the US or ARIA in the UK) absorbs the risk of that early, seemingly impossible phase, it creates the breakthroughs the private sector eventually scales.
Philanthropic capital can move faster than government and tolerate more ambiguity than most venture firms. It’s best deployed at work where the commercial pathway isn’t yet legible, too applied for basic research grants but too early-stage and too uncertain for venture. Validation platforms, wet lab partnerships with AI companies, challenge prizes that define the problem before anyone can solve it. This is where I think Google.org and other philanthropic contributions are most differentiated.
Most venture capital becomes meaningful once the hypothesis is reasonably clear and the question is about execution. It's less well-suited to the stage when you're still figuring out what the right question is.
But the most interesting work in this space is increasingly happening at the boundaries between those layers, where the field lacks effective infrastructure to drive progress. Consider AI-accelerated drug discovery for neglected tropical diseases, where the commercial market is too small for pure venture but the problem is too applied for a research grant. Or open foundation models for scientific domains, where building them is a public good but the infrastructure around them has real commercial dimensions. Recoverable grants, revenue-based structures, and mission-aligned LP arrangements are the tools for exactly these situations, and building them out is one of the most important things this strategy is trying to accelerate.
TK: The strategy is explicit that the gap between a promising AI prediction and an experimentally verified result is one of the most underfunded problems in the field. What does it actually take to bridge that gap?
DC: The Valley of Death is real, but it’s worth being precise about what kind of problem it is. It isn’t simply a funding gap. It’s a reality gap. An AI model can predict, with high confidence, which molecule might cure a disease or which material might store energy more efficiently. But that prediction exists only on a screen. Turning it into something you can test, touch, or treat with requires wet labs, physical materials, and experiments that can cost hundreds of thousands of dollars each. And as science agents enter the picture, the stakes rise further: any hallucination upstream doesn’t stay upstream. It propagates into experimental design, resource allocation, and potentially clinical decisions. That’s the distance we’re talking about.
What it takes to bridge that gap is genuine partnership between computational and experimental scientists, built before the prediction exists. Not assembled as an afterthought once the results arrive. That doesn’t mean constraining what AI can find. Unconstrained search sometimes surfaces things no human researcher would look for, and that’s part of the value. But the experimental partnership needs to be in place before the results arrive, not scrambled together afterward.
The harder observation is that incentive structures in both academia and parts of the AI industry actively reward staying in simulation. That’s where you control the variables, where papers get written, where demos look clean. Going into the real world means accepting friction, cost, and negative results. The models that work have made a cultural decision to do it anyway: embedded experimental capacity, funding structures that cover validation costs without requiring them to be pre-justified, and a genuine tolerance for finding out you were wrong.
Tolerating negative results isn’t just cultural, it’s infrastructural. We treat negative results as waste, when they’re actually signal. A shared database of them could be one of the most valuable contributions this field makes, not just capturing outcomes, but preserving what the experimental process itself revealed: the conditions that failed, the assumptions that didn’t hold, the unexpected behavior along the way. That’s where a lot of the real knowledge lives.
The Wellcome Leap programs are one example of getting this right. Some ARPA-H program structures are attempting to build it in by design. The Cancer Research UK Grand Challenge model has real elements of it. What they share is a willingness to fund the friction, not just the prediction.
TK: The strategy calls for identifying a focused set of "grand challenges" rather than spreading effort across the whole field. How do you go about selecting those? What does a genuinely AI-shaped scientific problem look like in practice?
DC: The selection process for grand challenges is more iterative than people expect. You don’t identify them simply by looking at the scientific literature and picking the hardest problems. You identify them by looking for the intersection of three things: a problem that is genuinely AI-shaped, meaning that scale, pattern recognition, or hypothesis generation is the binding constraint; a community of scientists who are ready to work at the boundary of their discipline and willing to share data; and a validation pathway that is difficult but not impossibly expensive.
The problems that look most promising are those where large amounts of underutilized data can be collated, structured, and made ML-ready; where lab automation has brought validation costs down; and where the potential impact is vastly disproportionate to what the field is currently spending. Parts of infectious disease, metabolic disorders, materials discovery, and climate adaptation science meet those criteria.
For researchers: the test I find most useful is whether AI is doing something structurally impossible at human scale, not just faster or cheaper, but categorically different. If the answer to that question is yes, you have something worth bringing.