AI Coding & Data Agents for Researchers

v2.2.0 · updated 2026-09-11
A privacy-first guide to AI coding and data-analysis tools for scientific work. Start from your data, see what's appropriate, and what it costs in capability. As open as possible, as closed as necessary.
Version 2.2.0: current first-party public summaries reviewed on 2026-09-11. Route-specific limits, access restrictions and held candidates are documented in the September live audit; verification is not institutional or contractual approval. Privacy and pricing change fast — confirm your exact plan/tier and region before relying on any row.
Classify your data first (ELIXIR RDMkit · Data sensitivity). This is decision-support, not legal advice — confirm with your DPO / data steward and the vendor's own terms before using sensitive or regulated data. Privacy & pricing change often; each row shows when it was last checked.
What data will you put into the tool?
Non-sensitive = public / anonymised / synthetic  ·  Personal = pseudonymised (GDPR)  ·  Special-category = health / genetic / clinical (GDPR Art. 9)
Pick your data class to see what's suitable — or browse all below.
What the data-handling labels mean — Local supports local/self-hosted inference; configure providers, telemetry and integrations Zero-retention documented content-retention limits for a qualifying route; exclusions and product storage can differ No-train documented no-training scope; plan, provider and storage exceptions can differ Opt-out trains on your data unless you turn it off Trains by default your inputs train the vendor's model Unclear training/retention not established for this route; verify current terms
Type
Data handling
Capability
Only
⚡ Best starting points — by need
ToolTypeData handlingCapability Non-sens. Personal Special Links