First call on compute() on some data field, e.g. as in
import scida
import dask.array as da
import numpy as np
from time import time
TNG_MODEL = "50"
sim = scida.load(f"TNG{TNG_MODEL}-2")
ds = sim.get_dataset(z=3.0)
h = 26
data = ds.return_data(haloID=h)
bh = data["PartType5"]
mass = bh["BH_Mass"].magnitude
%time mass.compute()
is very slow:
CPU times: user 9.46 ms, sys: 63.6 ms, total: 73.1 ms
Wall time: 7.43 s
opposed to subsequent calls
CPU times: user 6.75 ms, sys: 22.4 ms, total: 29.1 ms
Wall time: 71.9 ms
Turns out this linked to
1.) MPCDF /virgotng filesystem where the simulation is stored has unusually long first file access times.
2.) Even for slices from a single file, all files contributing to a given virtual dataset are opened, which (when combined with previous problem) leads to the long loading times. Latter behavior becomes apparent when running lsof -a -p $(pgrep python | tail -n1) | egrep 'REG' to query for the python session's file handles.
Reading a single entry in above example will open all underlying files
cbyrohl@workstation-cbyrohl:~/repos/scida$ lsof -a -p $(pgrep python | tail -n1) | egrep 'REG'
python 3074680 cbyrohl 57rR REG 0,38 2907815800 2365075 /newdata/data/public/testdata-astrodask/TNG50-2_snapshot_z3/snap_025.0.hdf5
[...]
python 3074680 cbyrohl 184rR REG 0,38 2981622844 2361001 /newdata/data/public/testdata-astrodask/TNG50-2_snapshot_z3/snap_025.127.hdf5
while directly reading the cachefile that contains the virtual dataset via
hf = h5py.File("/home/cbyrohl/cachedir/7b10575642ef0f4e/data.hdf5")
m = hf['PartType0']["Masses"]
print(m[5]) # only first file is opened
m2 = da.from_array(m)
m2[5].compute()
just opens the required file
cbyrohl@workstation-cbyrohl:~/repos/scida$ lsof -a -p $(pgrep python | tail -n1) | egrep 'REG'
python 3073903 cbyrohl 50rR REG 0,38 2907815800 2365075 /newdata/data/public/testdata-astrodask/TNG50-2_snapshot_z3/snap_025.0.hdf5
Hence, it appears possible for hdf5+zarr to only open/read the needed file as expected, but somehow this does not appear to be the case for the scida wrapper.
First call on compute() on some data field, e.g. as in
is very slow:
opposed to subsequent calls
Turns out this linked to
1.) MPCDF /virgotng filesystem where the simulation is stored has unusually long first file access times.
2.) Even for slices from a single file, all files contributing to a given virtual dataset are opened, which (when combined with previous problem) leads to the long loading times. Latter behavior becomes apparent when running
lsof -a -p $(pgrep python | tail -n1) | egrep 'REG'to query for the python session's file handles.Reading a single entry in above example will open all underlying files
while directly reading the cachefile that contains the virtual dataset via
just opens the required file
Hence, it appears possible for hdf5+zarr to only open/read the needed file as expected, but somehow this does not appear to be the case for the scida wrapper.