# Save large datasets (on unstructured grid) in PVD (.VTU) format

**URL:** https://discourse.paraview.org/t/save-large-datasets-on-unstructured-grid-in-pvd-vtu-format/1833
**Category:** ParaView Support
**Created:** [May 11, 2019, 12:54pm UTC](https://discourse.paraview.org/t/save-large-datasets-on-unstructured-grid-in-pvd-vtu-format/1833 "2019-05-11T12:54:42Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![PasCo31](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/p/9d8465/32.png) [@PasCo31](https://discourse.paraview.org/u/PasCo31)
#### Post date: [May 11, 2019, 12:54pm UTC](https://discourse.paraview.org/t/save-large-datasets-on-unstructured-grid-in-pvd-vtu-format/1833/1 "2019-05-11T12:54:42Z")

</div>

Hi everyone,

I’m using pvbatch to post-process DNS data and although calculations scripted in Python are quite fast, saving the data requires a prohibitively large amount of time.  
Here is the function used to save the data:  
SaveData(’/path/file.pvd’, proxy=activeSource )  
It’s also taking a long time to save them when launching pvbatch in parallel.

Is there a more time-efficient way to save such large datasets?

Thanks in advance for your help.

Best,  
-Pascal

---

<div class="post-metadata">

### Author: ![berkgeveci](https://discourse.paraview.org/user_avatar/discourse.paraview.org/berkgeveci/32/1464_2.png) [@berkgeveci](https://discourse.paraview.org/u/berkgeveci)
#### Post date: [May 14, 2019, 3:30pm UTC](https://discourse.paraview.org/t/save-large-datasets-on-unstructured-grid-in-pvd-vtu-format/1833/2 "2019-05-14T15:30:01Z")

</div>

Can you elaborate on your workflow? Are you loading simulation output, filtering it and then writing out filtered data?

---

<div class="post-metadata">

### Author: ![PasCo31](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/p/9d8465/32.png) [@PasCo31](https://discourse.paraview.org/u/PasCo31)
#### Post date: [May 14, 2019, 5:47pm UTC](https://discourse.paraview.org/t/save-large-datasets-on-unstructured-grid-in-pvd-vtu-format/1833/3 "2019-05-14T17:47:20Z")

</div>

Here’s the workflow (steps with \* are optional):

- Data loading from NEK5000
- Calculation of various quantities (on unstructured grid with around 50e6 points and circa 400 time instances)  
-\* Gaussian filtering
- Temporal statistics  
-\* Slicing
- Saving data (on unstructured grid, structured grid or slice) --\> sticking point = very high amount of time (days)

I also noticed that the reader for NEK5000 fields isn’t stable. It seems that sometimes (for several time instants) the indexing of the points isn’t done correctly (see attached figure).

 ![ProblemParaviewNEK](https://discourse.paraview.org/uploads/default/original/2X/9/9c40c737589afff59fdc0611873607f1eb55e029.png)

Many thanks.

Best,  
-Pascal

---

<div class="post-metadata">

### Author: ![berkgeveci](https://discourse.paraview.org/user_avatar/discourse.paraview.org/berkgeveci/32/1464_2.png) [@berkgeveci](https://discourse.paraview.org/u/berkgeveci)
#### Post date: [May 16, 2019, 4:51pm UTC](https://discourse.paraview.org/t/save-large-datasets-on-unstructured-grid-in-pvd-vtu-format/1833/4 "2019-05-16T16:51:21Z")

</div>

For sanity check, can you try the legacy VTK format and the Exodus format for your unstructured grid use case? I’d like to know how they perform.

---

<div class="post-metadata">

### Author: ![PasCo31](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/p/9d8465/32.png) [@PasCo31](https://discourse.paraview.org/u/PasCo31)
#### Post date: [May 20, 2019, 12:12pm UTC](https://discourse.paraview.org/t/save-large-datasets-on-unstructured-grid-in-pvd-vtu-format/1833/5 "2019-05-20T12:12:04Z")

</div>

Using legacy VTK and Exodus formats does not speed the process up.  
Furthermore, when I apply the temporal statistics filter to the unfiltered DNS fields, I obtain an averaged field with misplaced points like the ones in the figure previously sent.

Attached you will find two examples of Python scripts used to run the post-processing steps with pvbatch on our cluster.  
[macroPostProcDNSF\_5000\_2\_tot.py](https://discourse.paraview.org/uploads/default/original/2X/a/a3722e87cf4962696b7255abadcb55be63eb621d.py) (19.6 KB)  
[macroPostProcDNSMeanHighRes2.py](https://discourse.paraview.org/uploads/default/original/2X/9/9ca44f812a9d9e744ee51fef0743ed68e67741b5.py) (9.7 KB)

Many thanks.

---

<div class="post-metadata">

### Author: ![juanecopro](https://discourse.paraview.org/user_avatar/discourse.paraview.org/juanecopro/32/2285_2.png) [@juanecopro](https://discourse.paraview.org/u/juanecopro)
#### Post date: [October 23, 2019, 4:43pm UTC](https://discourse.paraview.org/t/save-large-datasets-on-unstructured-grid-in-pvd-vtu-format/1833/6 "2019-10-23T16:43:36Z")

</div>

> [@PasCo31](#):
>
> sing legacy VTK and Exodus formats does not speed the process up.  
> Furthermore, when I apply the temporal statistics filter to the unfiltered DNS fields, I obtain an averaged field with misplaced points like the ones in the figure previously sent.
> 
> Attached you will find two examples of Python scripts used to run the post-processing steps with pvbatch on our cluster.

Hi Pascal,  
I have a quick comment. I believe the Nek5000 reader for Paraview depends on the VisIt reader. I know that in VisIt the element numbering can be messed up because it bases the mesh on the first file in your batch, i.e. if you have binary files fld0.f%05d, the reader takes the mesh from file fld0.f00001. If you obtained the files from different runs, the meshes might be inconsistent. At least this is an issue I’ve ran into quite a few times, specially when I create new files from a post-processing step, for example, when computing vorticity from dumped velocity data. In that case it is useful to dump out grid data (ifxyo = .true. in Nek5000) for at least the first file of the sequence. Just keep this in mind.

Juan Diego
