home / scout

scout

apt-scout is built around a hand-curated, verbatim-verified catalogue of aptamer–target binding affinities (Kd) — not a re-hosted aggregation. Every Kd is read back from its primary publication, stored with the exact source sentence, and labelled by measurement type (intrinsic equilibrium vs apparent/avidity); a verbatim re-verification found that ~40–50% of reported Kd values in existing aggregated aptamer collections are in fact apparent/avidity, not comparable. This curated Kd layer is then linked to a target-context profile of ~7,177 human proteins (the layers below are harvested from public databases and labelled as such).

apt-scout is an automated catalogue of 7,177 human protein targets, each evaluated across a five-layer schema (structure / modification / interaction / expression / disease) to prioritise targets for affinity-reagent (aptamer, antibody, ASO, small-molecule) design — with explicit annotation of activation-state-selective surfaces.

How the data was built (provenance): values are mechanically harvested from public databases — PDB / AlphaFold / EMDB (structure), UniProt (modifications), Open Targets / ChEMBL / PubMed / Aptagen / HPA (interaction), Human Protein Atlas / Vesiclepedia / ExoCarta / EV-Map (expression), Open Targets (disease). The raw source response is stored per target in each layer's raw_json column, so every value is traceable to its origin. Activation-state PDB pairs are a human-curated list of real PDB structures. No free-text was generated by an LLM in these layers.

How to verify any value: open a target, then check its source — e.g. PDB IDs at rcsb.org, UniProt at uniprot.org, expression at proteinatlas.org, disease at platform.opentargets.org. The raw_json column shows the exact upstream record (including PubMed queries + PMIDs).

Tiers: Tier 1 = 3,300 PDB-anchored targets; Tier 1.5 = 3,874 AlphaFold-confident targets; plus 3 controls.

Completeness / known limitations (transparent): some fields are not yet fully populated — ppi_count is pending; OMIM / ClinVar counts are blank in this production snapshot (held in a separate reference cohort); AlphaFold pLDDT is missing for 653 targets. A data-quality audit (incl. inter-rater verification) is in progress.

JSON API: append .json to any table or query. License: CC BY 4.0.

Data license: CC BY 4.0 · Data source: apt-scout automated curation pipeline (E. Dohi, NCNP) — values harvested from public databases; raw source stored per target

Custom SQL query returning 23 rows (hide)

This data as json, CSV

rowidordmetricvaluegradedefinitionsource
1 1 human targets (published catalogue) 7177 defined-set Human protein targets in the PUBLISHED CATALOGUE = harvested across five evidence layers (Tier 1 PDB-anchored + Tier 1.5 AlphaFold + controls). 'Published' here means included & harvested, NOT that all are deep-curated. See the curation-stage ladder (#20-21): registry -> published -> scored. Deep LLM 5-layer curation is a smaller, growing backend subset. apt-scout target registry
2 2 with experimental 3D structure (PDB) 4428 harvested Targets with at least one experimental PDB structure. RCSB PDB (layer 1)
3 3 with cryo-EM structure 1683 harvested Targets with at least one cryo-EM structure. RCSB PDB / EMDB
4 4 AlphaFold-predicted only (Tier 1.5) 3874 predicted Targets with no experimental PDB; AlphaFold model only (predicted, not experimental). Completed targets only, so Tier 1 (3,300) + Tier 1.5 (3,874) + 3 controls = 7,177. AlphaFold DB
5 5 curated active/inactive PDB pairs 11 human-curated Targets with a hand-curated active vs inactive PDB pair. manual curation (layer 2)
6 6 aptamer/SELEX literature hits 1472 keyword-unverified PubMed '(gene) AND (aptamer OR SELEX)' returned a hit. Co-mention only; includes false positives; NOT a verified aptamer. PubMed keyword (layer 3)
7 7 detected in EV-Map plasma dataset 3422 detected apt-scout targets detected in the EV-Map plasma-EV proteome (broad detected set; NOT the conserved proteome). Rai & Greening 2025 Nat Cell Biol
8 8 EV-Map conserved EV-hallmark (in apt-scout) 104 source-derived Of the EV-Map 182 conserved EV-hallmark proteins, those that are apt-scout targets (gene-matched to source supp7). Rai & Greening 2025 supp7
9 9 verbatim-verified Kd measurements 742 verbatim-verified Distinct aptamer-target Kd measurements, each with a verbatim source quote. corpus literature extraction (v4)
10 10 intrinsic-equilibrium Kd (comparable) 555 verbatim-verified Kd measurements that are intrinsic equilibrium (the comparable subset for ML). corpus (v4)
11 11 verified aptamers (with Kd) 560 verbatim-verified Distinct aptamers with a verified, quantified Kd (contrast with the 1,472 keyword hits). corpus (v4)
12 12 targets in Kd layer 300 verbatim-verified Distinct targets (proteins/glycans/cells) in the binding-affinity layer. corpus (v4)
13 13 cell-surface / ecto targets (PREDICTED EV-surface accessible) 1136 predicted Integral membrane / ecto-domain proteins (CD markers, GPCRs, ion channels, integral membrane). EV biogenesis normally preserves topology so the ectodomain is PREDICTED to face the EV surface. NOT a measurement: HPA reports cellular (not EV) localization; lipid asymmetry can partially flip (PS via scramblase on activated/platelet EVs); cytoplasmic proteins can attach as a corona. Confirm by protease-protection / intact-EV surface labelling / immuno-EM. HPA subcellular + protein class (Thul 2017 / Uhlén 2015)
14 14 EV-surface aptamer candidates (surface ecto AND EV-Map hallmark) 24 predicted Top design set: cell-surface ecto proteins that are ALSO EV-Map conserved EV-hallmark proteins (detected on circulating EVs). PREDICTED accessibility (see #13 caveats) AND-ed with EV-Map detection evidence. Ranked in v_surface_targets. HPA + Rai & Greening 2025
15 15 Kd records human-verified (stratified sample) 0 human-curated Kd records with a LOGGED human verdict (confirmed/corrected) from the ongoing stratified-random verification. Grows post-publication; the rest are multi-agent / extraction verified. This is the honest, version-tracked QC status. human verification ledger
16 16 Kd records multi-agent verified (L2) 181 source-derived Kd records that passed independent multi-agent (L2) adversarial verification but are not yet in the human sample. corpus L2 pipeline
17 17 Kd records with a verbatim-verified sequence 435 verbatim-verified Kd records whose aptamer sequence is verbatim-verified against the source text/SI (clean ACGTU; modifications in the chemistry columns). The rest are flagged sequence_status=pending (sequence only in a figure or paywalled SI), being curated post-submission. corpus sequence backfill v1 (2026-06-22)
18 18 aptamer-protein co-structures (binding-site precedent) 114 harvested Experimental aptamer-protein co-structures (PDB): demonstrated cases where an aptamer binds a protein, with the binding location known. Empirical aptamer-amenability evidence. RCSB PDB (corpus structure handoff)
19 19 targets with BOTH measured Kd and a co-structure 33 source-derived Targets where affinity (Kd) AND binding location (co-structure) are both known — the highest-value set for structure-guided, modification-aware design. See kd_structure_crossmatch. corpus Kd x PDB crossmatch
20 20 target registry (all rows, incl. queued/failed) 7198 defined-set Every target row in the registry, including those not in the published catalogue. Registry -> published (#1) -> scored (#21). apt-scout target registry
21 21 targets with a heuristic priority score 6160 heuristic Published targets that carry a heuristic_priority_score (the rest are unscored). The score is a model-assisted ranking AID, not a validation — its inputs are LLM-estimated (see the v_targets column note). apt-scout scoring (heuristic)
22 22 Kd records with a resolved DOI 742 source-derived Gold Kd records whose source publication DOI was resolved from its PMID (NCBI E-utilities). Complements the always-present PMID + verbatim quote. NCBI E-utilities (PMID->DOI)
23 23 Kd records with a UniProt target id 552 source-derived Gold Kd records whose protein target carries a UniProt accession (reconciled v4+260613 map, non-destructive). The remainder are small molecules / organisms / complexes with no single UniProt entry. apt-scout UniProt reconcile
Powered by Datasette · Queries took 7.07ms · Data license: CC BY 4.0 · Data source: apt-scout automated curation pipeline (E. Dohi, NCNP) — values harvested from public databases; raw source stored per target