home

Home · About · Data sources · How to cite · Help & download · Numbers explained

apt‑scout

A hand‑curated, measurement‑class‑aware, verbatim‑anchored catalogue of aptamer–target binding affinities (Kd) — linked to a sortable profile of 7,177 human protein targets (structure, modifications, interactions, expression, disease). The linkage is the point: the curated Kd layer is cross‑linked to protein‑target context — including 114 experimental aptamer–protein co‑structures that pin the binding site — so affinities can be read in structural context. (EV, disease and druggability are added, secondary layers.) Every Kd is curated from its primary publication, carries the exact source sentence, and is labelled by measurement type (intrinsic vs apparent/avidity) so values are actually comparable. Use it to choose & prioritise targets for aptamer / antibody / ASO / small‑molecule reagents.

742curated, verbatim‑verified Kd measurements
555intrinsic‑equilibrium (directly comparable)
300targets · 287 publications
435with a verbatim‑verified sequence (59%)
7,177linked human protein targets (context)
Try: ITGB3 ABL1 HRAS Browse & sort the full table →

Explore the table — sort & filter

The main view is one sortable, filterable table with gene names. Click any column header to sort (e.g. most PDB structures, highest disease score). Use the facets (right side) to filter by tier, known aptamer, EV‑Map membership, activation-state pair, or cryo-EM. Add .json to any URL for the API.

Quick filters: Aptamer/SELEX literature hits (1,472, PubMed keyword — unverified) In EV-Map dataset (3,422 detected) EV-Map conserved proteome (182+42) EV & aptamer (843) Activation pair (11)

Binding affinities (Kd) — curated, not aggregated

The distinctive layer — and not a re-hosted aggregation. Each Kd was read back from its primary publication, stored with the exact source sentence (verbatim_quote) and classified by measurement type. Why it matters: in a verbatim re-verification, ~40–50% of reported "Kd" values in existing aggregated aptamer collections turned out to be apparent / avidity measurements, not intrinsic equilibrium constants — not comparable, and silently corrupting any model trained on the pooled values. apt-scout separates them so you can filter to a genuinely comparable set.

The core of apt-scout: aptamer–target binding affinities (Kd) matched to each target's structure and evidence, so you can see which proteins already have a quantified aptamer — and, by sorting/filtering, which do not (the reagent gaps). Intrinsic equilibrium Kd (vs purified target) is kept separate from apparent / avidity measurements, which differ by orders of magnitude. Sort by kd_log10_molar (lower = tighter).

Quick views: Intrinsic Kd, tightest first → Browse all 742 Kd measurements Intrinsic only (555) Benchmark / reference subset (comparable Kd) →

How to use: values are verbatim from papers and units are mixed — do not use the raw reported values directly; sort/compare/train on kd_log10_molar and filter to intrinsic + protein (or grab the ready-to-use subset). Source-traceable Kd layer (measurement-class-labelled, verbatim-anchored; human-verified n=0 so far, verification in progress): 742 distinct measurements (1 review-only value excluded on manual review), 555 intrinsic, 300 targets, 287 publications; 435 (59%) carry a verbatim-verified aptamer sequence (the rest flagged sequence_status=pending — sequence reported only in a figure or paywalled SI, being curated post-submission). Targets include proteins, glycans and small molecules. Source: corpus literature-extraction pipeline (E. Doi).

Aptamer amenability & Kd ↔ structure linkage

Not every protein is equally aptamer-tractable. apt-scout links the curated Kd layer to experimental aptamer–protein co-structures, so you can see, per target, whether an aptamer is even known to bind (a binding-site precedent) and where.

Quick views: Targets with BOTH Kd & co-structure (33) → All aptamer–protein co-structures (114)

Amenability caveat (design-relevant). Aptamers have an anionic backbone, so they strongly favour basic / positively-charged protein interfaces. In a preliminary analysis of the 114 co-structures, of 111 computable interfaces ~103 were positive, 5 neutral, only 3 negative — and those 3 bound by shape/hydrophobicity, not charge complementarity. Acidic / negative epitopes (e.g. the activation-specific surface of integrin αIIbβ3) have essentially no aptamer precedent and are better addressed by small-molecule or chemically-modified-aptamer approaches. (Preliminary interface-charge analysis, E. Dohi 2026; per-target epitope-charge scoring is a planned structural layer.)

On the EV surface or inside as cargo? — predicted membrane topology

An optional lens for EV work. An aptamer on an intact extracellular vesicle can only bind an extracellular-facing epitope. EV biogenesis normally preserves membrane topology (the ectodomain faces outward, the cytoplasmic tail faces the lumen), so we predict EV-surface accessibility from each protein's cell topology (Human Protein Atlas subcellular localization + protein class): integral / ecto cell-surface proteins (CD markers, GPCRs, ion channels) vs plasma-membrane peripheral proteins on the cytoplasmic leaflet (SRC, LYN, RHOA…) vs secreted / surface-corona vs luminal cargo.

Quick views: EV-surface candidates (surface × EV-Map hallmark, ranked) → All cell-surface ecto targets (1,136) Surface vs cargo counts

1,136 integral/ecto cell-surface · 486 PM-peripheral (cytoplasmic leaflet) · 539 secreted/corona · 4,184 luminal cargo · 852 no HPA localization. Source: Thul et al. 2017 & Uhlén et al. 2015 (HPA).

This is a prediction, not a measurement — please confirm experimentally. Caveats: (1) HPA reports cellular localization, not EV localization; (2) lipid asymmetry can partially flip — phosphatidylserine externalizes via scramblase (TMEM16F) on activated / platelet / apoptotic EVs; (3) normally-cytoplasmic proteins can attach to the outer surface as a protein corona; EV populations are heterogeneous. Confirm accessibility by protease-protection (proteinase K ± detergent), intact-EV surface labelling / bead-capture flow, or immuno-EM. The targetability_score is a transparent composite, not an experimental value.

The five evidence layers (per target)

1 · StructureExperimental PDB for 4,428 targets (1,683 cryo-EM); the rest are AlphaFold-predicted (Tier 1.5). Active/inactive pairs (PDB, AlphaFold, EMDB)
2 · ModificationPTM counts; activation-state structure pairs (UniProt; curated PDB list)
3 · InteractionDrugs, known aptamers & antibodies (Open Targets, ChEMBL, PubMed, Aptagen, HPA)
4 · ExpressionTissue expression; EV-proteome membership (Human Protein Atlas, EV-Map)
5 · DiseaseDisease associations & scores (Open Targets)

Download & API

Everything is freely downloadable — no login. Append .csv or .json to any table/query URL (add ?_stream=on&_size=max for all rows).

K‑d records (CSV) Targets (CSV) Kd (JSON) Full API & SQL help →

Trust & verify

All values are mechanically harvested from public databases (human proteins only) and the raw source record is stored per target — check anything in two clicks via the UniProt / PDB links. Two caveats: (1) "aptamer literature hits" (1,472) are PubMed keyword co-mentions, NOT verified aptamers — the verified, Kd-quantified aptamers live in the Binding-affinities layer (560 aptamers / 300 targets); (2) the priority_score is a heuristic ranking aid (0.30 biology + 0.30 reagent-gap + 0.20 druggability + 0.10 disease + 0.10 novelty) whose sub-scores are LLM-estimated — so it is not reproducible from fixed inputs and is not a validation or an experimental value; use it only to sort / short-list.

Human verification worksheet → CSV About How to cite Numbers explained (every figure + grade) → All tables & SQL

Transparent limitations: only 11 curated activation pairs; ppi_count & Vesiclepedia/ExoCarta not yet populated; OMIM/ClinVar blank in this snapshot; pLDDT missing for 653; headline figures cross-checked against the live database (see Numbers explained). CC BY 4.0 · E. Dohi, NCNP.

Data integrity & control. The public site is strictly read-only — visitors may browse, filter and run SELECT queries (SQL is SELECT-only; all writes are rejected) and download everything, but cannot modify any value. Data is changed only by the maintainer, by re-deriving the database from a frozen canonical snapshot through a scripted, provenance-stamped build and re-deploying it; the served database is checksum-verified to match that build. There is no public login or admin panel by design.

Versioning. apt-scout is published live and growing: each release is frozen and (from submission) given a versioned DOI, so every reported figure stays reproducible as the database expands. See the current version & release metadata. Verification (human + multi-agent) continues across releases and is tracked per record (verification_level).

Powered by Datasette · Data license: CC BY 4.0 · Data source: apt-scout automated curation pipeline (E. Dohi, NCNP) — values harvested from public databases; raw source stored per target