home

← Home · About · Data sources · How to cite · Help

About apt-scout

apt-scout is a curated, source-verified resource for choosing and characterising affinity reagents (aptamers, antibodies, ASOs, small molecules) against human protein targets. It is built on two layers joined by target, and organised on two orthogonal axes (confidence tier × measurement class).

Two layers

Target layer~7,177 human proteins profiled across five evidence layers — structure, modification, interaction, expression, disease — each value mechanically harvested from a public database with the raw source record stored per target.
Binding-affinity (Kd) layerSource-verified aptamer–target Kd measurements. Every Gold record carries the exact source sentence (verbatim quote) and a measurement-class label, so values can be audited and filtered to a comparable set.

What makes it different

1 · Verbatim provenance. Existing aptamer-affinity collections store a reported Kd without the sentence it came from, so per-record errors cannot be audited at scale. In apt-scout, every Gold Kd record carries its verbatim_quote — the exact source sentence or table cell.

2 · Measurement-class awareness. Intrinsic equilibrium Kd (aptamer vs purified target) is kept separate from apparent / cellular and avidity / multivalent measurements, which differ by orders of magnitude and are not comparable. Users can filter to a single comparable class — essential for any quantitative or machine-learning use.

3 · Reproducible pipeline. The public database is re-derived deterministically from a frozen, content-hashed snapshot plus scripted build steps, so it is reproducibly extendable as new papers appear (not a manually frozen dump).

Honest limitations (transparency)

Target-layer values are mechanically harvested (traceable, but not yet independently re-verified end-to-end); an inter-rater verification (κ) is in progress. priority_score is a model-assisted heuristic, not an experimental value. Some target fields are not yet populated (ppi_count; Vesiclepedia/ExoCarta; OMIM/ClinVar blank in this snapshot; pLDDT missing for 653). The Gold Kd set is an early, growing release; target_uniprot is populated only where the target is a human protein. All of these are stated, not hidden.

See every data source → How to cite All tables & SQL

Freely available, no login or registration. License CC BY 4.0 · E. Dohi, NCNP.

Powered by Datasette