Skip to content

Performance issue and extensive file access on (first) data read #164

Description

@cbyrohl

First call on compute() on some data field, e.g. as in

import scida
import dask.array as da
import numpy as np
from time import time
TNG_MODEL = "50"
sim = scida.load(f"TNG{TNG_MODEL}-2")
ds = sim.get_dataset(z=3.0)
h = 26
data = ds.return_data(haloID=h)
bh = data["PartType5"]
mass = bh["BH_Mass"].magnitude
%time mass.compute()

is very slow:

CPU times: user 9.46 ms, sys: 63.6 ms, total: 73.1 ms
Wall time: 7.43 s

opposed to subsequent calls

CPU times: user 6.75 ms, sys: 22.4 ms, total: 29.1 ms
Wall time: 71.9 ms

Turns out this linked to

1.) MPCDF /virgotng filesystem where the simulation is stored has unusually long first file access times.
2.) Even for slices from a single file, all files contributing to a given virtual dataset are opened, which (when combined with previous problem) leads to the long loading times. Latter behavior becomes apparent when running lsof -a -p $(pgrep python | tail -n1) | egrep 'REG' to query for the python session's file handles.

Reading a single entry in above example will open all underlying files

cbyrohl@workstation-cbyrohl:~/repos/scida$ lsof -a -p $(pgrep python | tail -n1) | egrep 'REG'
python 3074680 cbyrohl 57rR REG 0,38 2907815800 2365075 /newdata/data/public/testdata-astrodask/TNG50-2_snapshot_z3/snap_025.0.hdf5
[...]
python 3074680 cbyrohl 184rR REG 0,38 2981622844 2361001 /newdata/data/public/testdata-astrodask/TNG50-2_snapshot_z3/snap_025.127.hdf5

while directly reading the cachefile that contains the virtual dataset via

hf = h5py.File("/home/cbyrohl/cachedir/7b10575642ef0f4e/data.hdf5")
m = hf['PartType0']["Masses"]
print(m[5]) # only first file is opened
m2 = da.from_array(m)
m2[5].compute()

just opens the required file

cbyrohl@workstation-cbyrohl:~/repos/scida$ lsof -a -p $(pgrep python | tail -n1) | egrep 'REG'
python 3073903 cbyrohl 50rR REG 0,38 2907815800 2365075 /newdata/data/public/testdata-astrodask/TNG50-2_snapshot_z3/snap_025.0.hdf5

Hence, it appears possible for hdf5+zarr to only open/read the needed file as expected, but somehow this does not appear to be the case for the scida wrapper.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions