A dataset of high-resolution snapshots of the viscous sublayer from direct numerical simulation of a turbulent boundary layer up to 𝑅𝑒θ = 2400
File(s) DiB2026.pdf (885.13 KB)
Accepted version
Author(s)
O'Connor, Joseph
Whalley, Richard
Wynn, Andrew
Laizet, Sylvain
Type
Journal Article
Abstract
This dataset comprises time-resolved 3D fluid field data (pressure and the three velocity components)
from the viscous sublayer of a canonical zero-pressure-gradient turbulent boundary layer. In total it
contains 16,384 snapshots, amounting to approximately 11.1 TiB of data (pre-compression). In addition to the snapshot data, the dataset also includes time-averaged turbulent statistics over the full boundary layer for the four primary quantities (pressure and velocity) together with the second-order velocity products, enabling validation and comparison with existing literature. To create the data, direct numerical simulations were performed with the high-order flow solver Incompact3d on the ARCHER2 UK national supercomputer. Following the simulation, the raw Incompact3d outputs were converted to Zarr v3 and uploaded to a remote object store, together with accompanying materials (metadata, example scripts, licence, and readme). Other than format conversion, no additional processing has been applied. The data are hosted on the Edinburgh International Data Facility (EIDF), which provides a graphical web interface via the Comprehensive Knowledge Archive Network (CKAN) interface. Given the size and structure of the dataset, programmatic access is expected to be most
convenient; accordingly, the EIDF also exposes an interface compatible with a subset of the Amazon Simple Storage Service (S3) REST API. Performance-aware storage choices were made to facilitate
efficient remote access. Chunk sizes were selected to optimise anticipated common access patterns. Where appropriate, sharding was applied to reduce the number of files and the load on the remote filesystem and compression has also been applied to reduce network traffic. Example Python scripts demonstrate end-to-end usage (opening the stores, plotting, unit conversion, and chunk-awaresampling for machine-learning pipelines), lowering the barrier to entry and serving as templates for custom analyses. The dataset will enable a broad range of research activities, including developing and testing turbulence theory, training and evaluating data-driven models, and validating experimental protocols and lower-fidelity computational fluid dynamics models.
from the viscous sublayer of a canonical zero-pressure-gradient turbulent boundary layer. In total it
contains 16,384 snapshots, amounting to approximately 11.1 TiB of data (pre-compression). In addition to the snapshot data, the dataset also includes time-averaged turbulent statistics over the full boundary layer for the four primary quantities (pressure and velocity) together with the second-order velocity products, enabling validation and comparison with existing literature. To create the data, direct numerical simulations were performed with the high-order flow solver Incompact3d on the ARCHER2 UK national supercomputer. Following the simulation, the raw Incompact3d outputs were converted to Zarr v3 and uploaded to a remote object store, together with accompanying materials (metadata, example scripts, licence, and readme). Other than format conversion, no additional processing has been applied. The data are hosted on the Edinburgh International Data Facility (EIDF), which provides a graphical web interface via the Comprehensive Knowledge Archive Network (CKAN) interface. Given the size and structure of the dataset, programmatic access is expected to be most
convenient; accordingly, the EIDF also exposes an interface compatible with a subset of the Amazon Simple Storage Service (S3) REST API. Performance-aware storage choices were made to facilitate
efficient remote access. Chunk sizes were selected to optimise anticipated common access patterns. Where appropriate, sharding was applied to reduce the number of files and the load on the remote filesystem and compression has also been applied to reduce network traffic. Example Python scripts demonstrate end-to-end usage (opening the stores, plotting, unit conversion, and chunk-awaresampling for machine-learning pipelines), lowering the barrier to entry and serving as templates for custom analyses. The dataset will enable a broad range of research activities, including developing and testing turbulence theory, training and evaluating data-driven models, and validating experimental protocols and lower-fidelity computational fluid dynamics models.
Date Issued
2026-06-01
Date Acceptance
2026-04-07
Citation
Data in Brief, 2026, 66
ISSN
2352-3409
Publisher
Elsevier
Journal / Book Title
Data in Brief
Volume
66
Copyright Statement
Copyright © 2026 Copyright Owner. This is the author’s accepted manuscript made available under a CC-BY licence in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy)
License URL
Publication Status
Published
Article Number
ARTN 112767
Date Publish Online
2026-04-13
