Capstone Project

Computational binder design as a reproducible evidence trail

Capstone mission

Given a target protein and an intended binding surface, develop and defend a computational binder-design hypothesis. The deliverable is not a claim that a binder works. It is a reproducible portfolio showing how you selected a target and epitope, generated and sequence-designed candidates, validated complexes, learned from failure, and chose the next candidate or experiment.

This capstone preserves the final project from the original in-person bootcamp. The live groups presented for about 20 minutes; self-paced learners can submit an equivalent written report, narrated slide deck, notebook, repository, or video.

ImportantRead these before spending substantial compute

Capstone contract

Core capstone Full campaign
Active effort 10–25 hours 20–40+ hours
Compute time Add 5–30 hours of generation, sequence design, validation, and queues Add 20–100+ hours for larger sweeps, seeds, methods, and repeated validation
Scope One target, one defined epitope, a teaching-scale campaign, several sequence variants, and at least two finalists Broader hotspot/parameter/potential comparison with more candidates and independent checks
Output Seven-artifact evidence portfolio and completed rubric Core portfolio plus deeper ablations, expanded validation, and presentation-ready synthesis
Finished when All seven artifacts exist; rubric is at least 13/20 with no category below 2 The evidence trail is reconstructable and further compute is driven by a new hypothesis, not missing documentation

Compute ranges are planning estimates, not guarantees. Target length, campaign size, GPU, queue policy, model, and failed jobs dominate wall-clock time. Start every stage with the smallest smoke test that could expose a bad path, chain, residue number, environment, or output assumption.

Two valid participation modes

Generation mode

Run the pipeline yourself and preserve exact inputs, versions, commands/configurations, seeds, logs, raw outputs, and hardware context.

Analysis mode

If GPU, account, model, license, or quota access is unavailable, use supplied or previously generated outputs. Preserve their source and reconstruct as much provenance as possible. You still complete target analysis, filtering, confidence/interface interpretation, failure analysis, and a conditional selection memo. Clearly distinguish evidence you generated from evidence you inherited.

Both modes use the same rubric. Analysis mode is not permission to hide missing provenance; it is a scientifically honest fallback when infrastructure is outside the learner’s control.

Watch → Read → Do → Check

  • Watch (optional): the two original workshop recordings below provide the broad tool landscape and a practical workflow walkthrough.
  • Read: this overview, your selected target page, the rubric, and the relevant tool lessons.
  • Do: produce the seven staged artifacts below, preserving failed and rejected outputs as well as finalists.
  • Check: score the portfolio with the rubric, compare the decision logic with the worked example, and revise the weakest category.

Original workshop recordings

Part 1 — Cutting-edge ML tools for protein design pipelines (optional context)
Part 2 — AI-driven protein design workflow: a practical example (optional context)

Choose a target

Choose a target whose biological surface you can explain and whose structure you can prepare confidently. Do not choose solely because a name is familiar.

Target Starting structure Design context First question to resolve
PD-L1 4ZQK Immune-checkpoint interface Which PD-1-facing residues define a compact, accessible epitope?
IFNAR2 3SE3 Cytokine-receptor surface Which chain, state, and native interface support the intended competition hypothesis?
IL-7Rα 3DI2 Interleukin-receptor surface How do chain numbering and the native ligand interface map to hotspots?
Bet v 1 4A88 Allergen surface Which exposed patch supports the intended blocking or recognition hypothesis?
TrkA receptor 1WWW NGF-receptor complex Use the TrkA–NGF complex; 1WWC is an NT-3–TrkC structure, not an alternative TrkA target
TEM-1 β-lactamase 1FQG Enzyme surface/active-site context Should a binder occlude a functional site or target a distinct exposed surface?
GM2 activator protein 1G13 Lipid-transport protein Does the apo structure represent the pocket state relevant to the design mission?
β-Glucosidase 2JIE Enzyme/ligand complex How should the trapped intermediate and active-site geometry influence surface choice?

Each target page provides structure context, candidate hotspots, preparation guidance, and an interactive view. Verify chain IDs and residue numbers in the exact file you will use; numbering from a paper, UniProt, a biological assembly, and a deposited PDB may differ.

<img src="../images/course-dashboard/capstone-workflow.png" alt="Capstone workflow from target selection through backbone generation, sequence design, validation, and final selection.">
<p class="lesson-visual-caption">Treat the project as a sequence of evidence gates. Do not scale the next stage until the current artifact is reconstructable and interpretable.</p>
<img src="../images/course-dashboard/binder-target-interface.png" alt="Binder design target interface with hotspot residues on the target and a designed binder approaching the surface.">
<p class="lesson-visual-caption">A biologically motivated surface and a compact hotspot set provide a more testable design mission than an unrestricted target surface.</p>

Project stages and required evidence

Stage 1 — Define the target and epitope

Inspect the biological assembly, chains, residue numbering, missing regions, ligands/cofactors, glycans, alternate states, and native partners. State the mechanism you hope a binder would test.

Exit artifact: annotated target image plus a concise brief containing accession, chain/range, intended surface, three to six justified hotspots, and caveats.

Stage 2 — Freeze a reproducible run configuration

Record input identifiers/checksums, preparation commands, tool releases/commits, environment, checkpoints, hardware, exact commands/configurations, seeds, output paths, and deviations from the course’s known-good stack.

Exit artifact: another learner can reconstruct the smoke test without guessing a hidden setting.

Stage 3 — Generate and filter backbones

Begin with one or two outputs. Confirm the input chain, contig/chain break, hotspots, length, output location, and structure integrity. Then scale to a teaching-size campaign. Show the full distribution—not only favorites—and apply a consistent first-pass filter for topology, target/hotspot contact, clashes, and diversity.

Exit artifact: labeled gallery/table of successful and failed backbones with keep/reject reasons and attrition counts.

Stage 4 — Design sequences

Generate several sequences per advanced backbone. Preserve the parent-backbone link and record model/settings, fixed positions, temperature/seeds, recovery, diversity, composition, and relevant sequence liabilities.

Exit artifact: compact candidate table linking every sequence to its parent backbone and settings.

Stage 5 — Validate the designed complex

Predict the binder–target complex and inspect per-chain fold confidence separately from interface confidence and relative-chain uncertainty. Evaluate PAE, interface geometry, hotspot recovery, clashes, consistency across predictions, and at least one independent check beyond a single headline score.

Exit artifact: comparable figures and metrics for at least two finalists, with explicit uncertainty.

Stage 6 — Diagnose a failure and iterate

Choose at least one failed or misleading result. Form a cause hypothesis, make one targeted change, and ask whether the next run supports that diagnosis. Changing many settings at once weakens the inference.

Exit artifact: before/after comparison connecting failure → diagnosis → change → evidence → revised conclusion.

Stage 7 — Make a conditional selection

Define selection criteria before choosing a finalist. Compare the candidate you would advance with credible alternatives. State trade-offs, remaining uncertainty, and the next computational or experimental test that could change the decision.

Exit artifact: one-page final selection memo or equivalent final section.

Completion and self-assessment

Your format may be a report, notebook, repository README, narrated deck, or video, but the seven artifacts must remain easy to find and connected to their underlying files.

Use the Capstone Evidence Guide & Rubric to score:

  1. reproducibility;
  2. structural and computational evidence;
  3. scientific judgment;
  4. iteration and learning; and
  5. communication.

The course completion standard is 13/20 or higher with no category below 2. A high course score means the computational case is clear and defensible. It does not predict experimental success.

If your reasoning is becoming score-driven, compare it with the worked selection example. The exemplar’s strongest candidate is not the one with the highest summary score; it is the one whose evidence best matches the biological mission while making the remaining defect explicit.

What to do next

  1. Download the evidence template.
  2. Choose a target and draft Stage 1 before running a model.
  3. Review the Stage 2 configuration with the rubric before scaling compute.
  4. Stop at each exit artifact; repair missing provenance or interpretation before continuing.
  5. Share the final portfolio with a peer if possible and ask them to identify one unsupported inference.