Get ready
Open the folder in Codex
Click a step to see a screenshot.
Screenshots show the earlier practice guide and ChatGPT Work interface. Use your separate Codex setup instructions for current controls.
What’s in this folder?
Working with your own data? Keep good, up-to-date backups of all your data and know how to restore them.
Step 01 15 min
Discover what’s there
Understand the files before choosing what to do with them.
What’s in this folder, and how do the files fit together?
- Ask for file types, sizes, row counts and columns.
- Look for identifiers, relationships and handover notes.
- Leave
revealuntil you want to check your interpretation.
Need a nudge?
Inspect headers and a few rows first. Use code for full counts; large files do not need to fit into the chat. Ask which columns identify one record and which files should be stacked or joined.
Check your result
Demographic and enrolment batches describe participants. Tests contain repeated observations: participant, test, session and trial. Different tests have different units. Check for a duplicate export and avoid joins that multiply rows.
Step 02 10 min
Make the folder make sense
Choose useful names and a clear organisation.
Suggest useful names and an organisation for these files. Show me the plan first.
- Review the proposed names, evidence and uncertainties.
- Agree the changes, then ask the agent to apply them.
- Record old and new paths, edits and file hashes.
Check your result
A rename changes the path, not the file’s contents. Record a reason and new hash for edits. Document how you handle duplicates and notes. The pipeline should find the files after any approved renames.
Step 03 30 min
Merge, then check
Agree the relationships before writing code.
Help me merge the participant information and test results. Explain the table relationships and checks before writing code.
- Check unique IDs and the meaning of one row in each table.
- Keep tests and units distinct; missing values are not zero.
- Keep incomplete participants visible and document exclusions.
Ready to write the script?
Write a repeatable Python script for that plan. Record every transformation and flag duplicates, missing values and unmatched IDs.
Review inputs, outputs and dependencies. Decide who runs the script; ask before installations or new permissions.
Check your result
For the starting pack, the master has 100,000 unique participant IDs. Flag the duplicate export, repeated test rows and unmatched test ID. Check one participant against the source. Keep an issue report. If you changed the mock data, explain how the expected results change.
Step 04 20 min
Run it again
Show that the result can be reproduced.
Record how the merged dataset was made, and show how another person can reproduce it.
- Save inputs, hashes, commands, requirements and decisions.
- Run the pipeline again and compare outputs.
- Explain any changed inputs, unresolved issues and manual checks.
Check your result
The same inputs and decisions should produce the same output contents and hashes. Keep changing timestamps in the log, separate from data outputs. Record what the agent wrote or ran, what you approved and what you checked.
Step 05 Continue afterwards
Document these steps
Leave a workflow someone else can understand and use.
Document our steps and set up a repeatable workflow that tracks every change. Include a run guide, an append-only change log and verification checks. Show me the plan first.
What should the log capture?
Old/new paths, input/output hashes, commands, environment, decisions and checks. Include failed attempts and corrections. Append entries; do not overwrite earlier history. Log manual changes too, and compare hashes to catch omissions.
Check your result
Start a new conversation and ask the agent to use the run guide. Authorise a rerun. Check output hashes and confirm the new log entry was added without erasing the history.
Take it further
After agreeing eligibility rules, compare descriptive changes by condition or site, or build a dashboard from checked summaries. State denominators and limitations. These invented data support no clinical, efficacy or population claims.
When you’re ready
Check your understanding
Compare your discoveries with the description in the folder.
Open the reveal
Find sample-research-folder/reveal/dataset-description.md. The data dictionary is beside it. Compare both with your inventory and merge plan.
Optional second exercise
Explore a conference folder
Prefer documents and event information? Try last week’s fictional conference folder, including its editable example website.
All people, organisations and event details are synthetic. Extract it separately and open that folder in Codex.
What’s in this folder? Suggest useful tasks we could do, show your evidence and let me choose.