The context
In this illustrative scenario, a team develops models using images, documents or structured data. Datasets evolve through cleaning, annotation and enrichment. Computing environments change from one project to another, while the corpora need to remain identifiable, accessible and documented over time so teams can return to them for further work.
The problem to solve
Local copies multiply, and relationships between source data, prepared versions and experiments become difficult to track. Some transfers are repeated unnecessarily. The team also needs to check whether access performance suits the workloads it intends to run from its computing environments.
The proposed mission
- 01
Define an organisation for corpora and outputs using version conventions, manifests and the metadata needed to connect each experiment to the specific data it uses.
- 02
Test S3 access from preparation tools and computing environments, then measure read performance, transfers and batch behaviour using representative data and workloads from the team's projects.
- 03
Gradually introduce the chosen structure, document access rights and agree rules with the team for retention, useful duplication and removal of files that are no longer needed.
Intended benefits
- Better identified corpora that make the data used for an experiment easier to find.
- Fewer scattered copies through a shared structure and explicit rules for managing data.
- Clearer preparation of processing workloads and data transfers to the relevant computing environments.
How to measure the outcome
- Volume retained per corpus
- Data preparation time
- Share of experiments with a manifest
The baseline and objectives are defined during scoping, then compared with pilot results.