NewThe academic-hospital career path, step by step, with AI at every stepDiscover ›

Research

Organising a research database

From raw file to versioned database: disease-specific core datasets, a pseudonymised index, quality checks, and a query page that shows only aggregates.

Fellow (CCA, AHU)Associate prof. (MCU-PH)Full prof. (PU-PH)CoworkClaude CodeArtifactLong-term project

The problem

Years of cohort spreadsheets, exam exports and thesis data collections, each in its own format. The question “how many patients do we have with this phenotype and this data item?” takes a week.

Step-by-step method

  1. A written pipeline: original → raw → harmonised → database → analysis → export, each stage in its own folder; nothing is modified in place.

  2. Core datasets by disease, each with its version (candidate, frozen) and a freeze procedure.

  3. A pseudonymised pivot index to de-duplicate patients present in several cohorts, kept on the local workstation.

  4. Quality-control rules and a decision log (why a given variable was recoded).

  5. Distinguish “measurement in the database” from “file available”: you know what is actually usable.

  6. A query page that handles only aggregates: population, criteria, required data, result as counts.

  7. Identifiable extraction is done locally, by a script, on the basis of a validated “extraction request”.

  8. A “locked drawer” for sensitive data: whatever must never enter AI tools is isolated and documented, and open questions about access to it are settled with the institution.

  9. Exports for collaborators: aggregated figures and tables, never individual-level data.

Deliverable

A versioned, documented database, and an answer in minutes to “how many patients for this study?”.

Safeguards

Health data: the regulatory framework applicable to research (declaration, reference methodology, patient information), storage authorised by the institution, no identifying data in an unauthorised service. Consult your institution’s clinical research department and DPO (data protection officer).

Starter prompt

Here is the list (column names only, no data) of my 4 cohort files. Propose a harmonised data dictionary, quality-control rules, and the pipeline folder structure. Do not open any patient data.

← Back to the triptych