Author

Contributing guest author
Shiva Patre Manager AI Systems – Ellison Institute

Structure-based drug discovery has entered a new phase. The Ellison Medical Institute (EMI), a Los Angeles-based non-profit advancing cancer research and translational science, is partnering with Oracle to develop an AI-powered data platform that helps researchers transform millions of computational protein designs into governed, searchable, and actionable scientific insights.

Modern generative design pipelines can produce thousands, and eventually millions, of candidate protein binders for a single target. These designs often include rich computational signals: predicted binding confidence, structural quality metrics such as ipTM, ipSAE, pLDDT, sequence information, and protein structure files such as PDB or mmCIF formats. While this opportunity is enormous, so is the operational challenge. By organizing and prioritizing these data at scale, the platform provides faster, more informed decisions about how they should advance into experimental validation.

For a drug discovery team, the question is no longer simply: Can we generate more designs? The more important question is: Which designs should we trust enough to test experimentally?

This is the challenge the EMI AI Data Platform is designed to solve. EMI’s de novo mini-binder discovery workflow generates large volumes of unstructured protein structure files alongside structured metadata such as sequences and confidence metrics. As these datasets continue to grow, researchers contend with fragmented design data, limited discoverability across runs, manual candidate selection, incomplete lineage tracking, and governing scientific assets over time.

Built on Oracle’s data and AI capabilities, the platform transforms raw protein design outputs into governed, searchable, and actionable discovery intelligence. By connecting computational predictions with experiential outcomes, researchers can more efficiently prioritize the most promising candidates and continuously improve future design cycles.

For a drug discovery team, the question is no longer simply: Can we generate more designs? The more important question is: Which designs should we trust enough to test experimentally?

This is the challenge the EMI AI Data Platform is designed to solve. EMI’s de novo mini-binder discovery workflow generates large volumes of unstructured protein structure files alongside structured metadata such as sequences and confidence metrics. As these datasets continue to grow, researchers contend with fragmented design data, limited discoverability across runs, manual candidate selection, incomplete lineage tracking, and governing scientific assets over time.

Built on Oracle’s data and AI capabilities, the platform transforms raw protein design outputs into governed, searchable, and actionable discovery intelligence. By connecting computational predictions with experiential outcomes, researchers can more efficiently prioritize the most promising candidates and continuously improve future design cycles.

The drug discovery use case: prioritizing viable binders

In target-specific compound or binder design, computational models generate many candidate molecules or mini-binders against a biological target. Each candidate may appear promising for different reasons. One design may demonstrate strong predicted binding affinity. Another may exhibit high structural confidence. Another may closely resemble previously successful binders in vector space.

The dashboard screenshots illustrate this challenge in practice. One view summarizes design quality across a run: total generated designs, passing filters, pLDDT distributions, ipSAE distributions, pDockQ-style quality signals, and pass/fail cohorts. Another view supports nearest-neighbor search, allowing a researcher to select one promising design and find structurally or computationally similar candidates. This matters because researchers are not just looking for the “top score.” They are looking for patterns: which designs cluster near successful candidates, which high-scoring designs may still be risky, and which overlooked candidates may deserve experimental validation.

The core scientific challenge is that computational scores are useful filters but incomplete predictors. A design may score well in silico and yet fail in vitro because the AI prediction was biased, the protein lacks stability, or the assay behaves differently than expected. Therefore, the objective is to create a repeatable feedback loop that connects computational predictions with experimental outcomes, enabling researchers to continuously refine future design strategies.

The screenshots and the video below demonstrate how researchers can search for, compare, and identify similar protein designs, helping prioritize the candidates most likely to succeed experimentally.

Find Similar Designs

Figure 1. Similarity-search view: researchers select a promising binder and identify related candidates, helping prioritize candidates using structural and computational context rather than a single score.

Dashboard

Figure 2. Design-quality dashboard: confidence-score distributions

Dashboard

Figure 3. Design-quality dashboard: Filter pass rates

Why an AI Data Platform matters?

Drug discovery becomes a data infrastructure problem when design volume outpaces human review. A spreadsheet-driven process may work for tens of designs. However, it breaks down when pipelines create thousands of PDBs, multiple runs, multiple targets, embeddings, assay labels, model versions, and derived quality metrics.

The AI Data Platform addresses this by creating a unified data fabric. Raw assets remain in object storage, while structured design metadata, file references, embeddings, and experimental labels become query able in Oracle Autonomous AI Database. The proposed Model Visualization Platform (MVP) data model includes tables for runs, designs, design files, target embeddings, design embeddings, and future experiment labels such as assay type, outcome label, KD, kon, and koff.

This architecture is important because it allows the team to ask higher-value questions:

Which designs passed heuristic filters across all runs?
Which failed designs resemble previous successes?
Which targets produce higher validation rates?
Which metrics are actually predictive for an in vitro success?
Which protein design runs generated redundant candidates?
Which designs should be archived, retained, or promoted?

Oracle’s toolset fits this use case because it brings data engineering, AI-native search, governance, model training, and researcher-facing applications together in a single environment: AI Data Platform handles ingestion, metadata cataloging, and lifecycle hooks; Oracle Autonomous AI Database powers analytics and joins; object storage holds raw PDB assets; vector embeddings enable similarity search; Oracle Analytics Cloud and Dashboards surface results; and OCI Data Science or AI Data Platform Spark compute runs the Machine Learning workflows.

Functional flow: From design generation to experimental learning

Design Flow

Figure 4. Data flow from design generation through metadata, similarity search, experimental validation, and model feedback—showing how each result contributes to the next discovery cycle

This flow changes the discovery process from a one-way pipeline into a learning system. Instead of simply generating designs, selecting a few manually, and losing context, every design becomes part of an evidence trail. The model prediction, structural metrics, embeddings, file location, selection decision, and experimental result can all be connected.

PDB and mmCIF files can remain in OCI Object Storage, while Oracle Autonomous AI Database stores the structured metadata, metrics, file URIs, and vector embeddings needed for analytics and similarity search. This avoids forcing every artifact into the database while preserving traceability by cataloging stable design identifiers and file references.

AI Data Platform notebooks and Spark compute support large-scale preprocessing, filtering, embedding generation, and dataset creation. OCI Data Science provides a managed Jupyter and GPU environment for deeper model training and inference workflows. The document specifically illustrates interoperability with OCI Data Science supporting notebook development, terminal and Git access, prebuilt environments, GPU scaling, model deployment, object storage integration, and transfer of development code between Data science and AI Data Platform notebooks and generated vectors with AI Data platform into database vector tables.

Where Oracle tools create practical value

Unified visibility. This is the topmost value. Researchers can see total designs, target-level coverage, pass/fail rates, metric distributions, and candidate cohorts in one interface rather than across disconnected CSVs, folders, and scripts. A key advantage a researcher looks for is end-to-end unified visibility across data, models, similarity search, dataset designs, data curation and protein visualizations with a single unified visual interface. This consolidation is also what makes the platform AI-ready: unifying designs, embeddings, metrics, and provenance into one queryable layer that models and agents can act on directly.

Model improvement. The second value is one of the key improvements that comes out of this exercise. Once in vitro results are captured in the same environment as computational predictions, EMI can train models that predict experimental success rather than merely computational attractiveness. This is the key scientific leap. The platform can learn which combinations of pLDDT, ipSAE, ipTM, pDockQ-style metrics, sequence embeddings, structural embeddings, target class, and design provenance are most predictive of real validation.

Similarity-driven discovery. With vector embeddings, the platform can answer questions such as “find designs similar to this sequence” or “show binders close to this PDB file.” The process notes that semantic search can use native vector capabilities in 26ai, with multiple embedding models of different dimensionality are available for comparison, fine tuning and re-embeddings yet stored in an embedded vector database.

Governance and lifecycle control. Scientific assets need retention rules, lineage, auditability, and lifecycle policies. Low-quality designs should be pruned or archived safely, while promising designs should remain discoverable. The services highlight the need for metadata-driven governance, centralized technical and scientific metadata, lifecycle hooks, archival policy, and audit tracking. This helps tracking down scientific metadata validations, comparisons across iterative runs and data engineering curations

Researcher accessibility. Researcher’s accessibility to pertinent information at the time and granularity level is a key metric that increases researcher productivity to discovery decisions. Dashboards, OAC visualizations, and conversational interfaces make the data usable by scientists who may not want to write SQL. A researcher could ask: “Show me the top 10 designs for EGFR with ipTM above 0.8 and ipSAE above 0.6?” or “Which passing designs are nearest to this candidate?” Oracle Database features such as Select AI, APEX, AI Data Platform Agent workflows, and custom vector processing and search as possible paths for natural-language and semantic interaction.

Conclusion: moving from score-driven to evidence-driven discovery

The goal is not simply to build another dashboard, it is to create an intelligent, evidence-driven foundation for AI-enabled drug discovery.

In a traditional workflow, computational scores rank candidates, researchers manually select a subset for testing, and experimental results often live in separate systems. Through the AI Data Platform model, each design is connected across its entire lifecycle, from generation and scoring to similarity search, visualization, experimental validation, and model refinement. This continuous feedback loop enables researchers to understand what makes a design successful beyond computational predictions alone.

By bringing computational and experimental data together in a governed environment, the platform helps researchers prioritize candidates with greater evidence, reduce wasted experimental cycles, uncover promising designs that otherwise might be overlooked, and continuously improve predictive models as new data becomes available.  

Together, Oracle and Ellison Medical Institute are building more than a data platform, they are creating a scalable foundation for AI-ready scientific discovery. By transforming complex protein design data into actionable intelligence, this approach has the potential to accelerate drug discovery, enable better scientific decisions, and shorten the path from breakthrough research to new therapies for patients.]

Author

Shiva Patre is a Manager, AI Systems in the Applied AI and Advanced Molecular Medicine research group at the Ellison Medical Institute (EMI). He leads the development of scalable AI infrastructure and cloud data platforms that power EMI’s generative AI drug discovery and digital pathology research. He enjoys partnering across scientific and engineering teams to translate complex research into production-ready systems. He believes the best infrastructure enables scientists to focus on discovery, not just technology.

For more information

To explore more about the Oracle AI Data Platform, check out these resources:-