AI Coding & Data Agents for Researchers

A privacy-first guide to AI coding and data-analysis tools for scientific work. Start from your data, see what's appropriate, and what it costs in capability. As open as possible, as closed as necessary.
Catalog version 2.0.1 modified 2026-07-14; all 43 records carry an explicit 2026-07-13 evidence-check date, with provisional or conditional claims labelled conservatively. Privacy and pricing change fast — confirm your exact plan/tier and region before relying on any row.
Classify your data first (ELIXIR RDMkit · Data sensitivity). This is decision-support, not legal advice — confirm with your DPO / data steward and the vendor's own terms before using sensitive or regulated data. Privacy & pricing change often; each row shows when it was last checked.
What data will you put into the tool?
Non-sensitive = public / anonymised / synthetic  ·  Personal = pseudonymised (GDPR)  ·  Special-category = health / genetic / clinical (GDPR Art. 9)
Pick your data class to see what's suitable — or browse all below.
What the data-handling labels mean — Local runs on your own hardware; data never leaves Zero-retention cloud, but nothing stored or trained on (usually enterprise/API) No-train not used for training, but may be retained Opt-out trains on your data unless you turn it off Trains by default your inputs train the vendor's model
Type
Data handling
Capability
Only
⚡ Best starting points — by need
ToolTypeData handlingCapability Non-sens. Personal Special Links