# Import giant(50GB) data file into Paraview

**URL:** https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860
**Category:** ParaView Support
**Created:** [November 7, 2018, 10:51am UTC](https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860 "2018-11-07T10:51:27Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![ZiweiWu](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/z/ea666f/32.png) [@ZiweiWu](https://discourse.paraview.org/u/ZiweiWu)
#### Post date: [November 7, 2018, 10:51am UTC](https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860/1 "2018-11-07T10:51:27Z")

</div>

Hi all,  
I tried to import a 50GB netCDF file into Paraview but always run into Memory error like this:

 ![image](https://discourse.paraview.org/uploads/default/original/1X/267237e96864e45005db4ad21a7e4ca5bad1eb7e.png)  
So I’m wondering if there’s a limit on the size of input file in Paraview or the memoryError happened just because my local memory was too small? And is there any way that I can run this on my local machine or it has to be HPC?

Thanks!

---

<div class="post-metadata">

### Author: ![cory.quammen](https://discourse.paraview.org/user_avatar/discourse.paraview.org/cory.quammen/32/11193_2.png) [@cory.quammen](https://discourse.paraview.org/u/cory.quammen)
#### Post date: [November 7, 2018, 2:20pm UTC](https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860/2 "2018-11-07T14:20:46Z")

</div>

Do you have more than 50GB of RAM? How many timesteps are in the file?

---

<div class="post-metadata">

### Author: ![banesullivan](https://discourse.paraview.org/user_avatar/discourse.paraview.org/banesullivan/32/11714_2.png) [@banesullivan](https://discourse.paraview.org/u/banesullivan)
#### Post date: [November 7, 2018, 3:30pm UTC](https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860/3 "2018-11-07T15:30:37Z")

</div>

I see this is coming from a `vtkPythonAlgorithm`. Is this a custom netCDF reader plugin?

If so, you may want to update the request data script to only read chunks of the netCDF file based on your current time step or data extent.

---

<div class="post-metadata">

### Author: ![ZiweiWu](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/z/ea666f/32.png) [@ZiweiWu](https://discourse.paraview.org/u/ZiweiWu)
#### Post date: [November 8, 2018, 3:15am UTC](https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860/4 "2018-11-08T03:15:24Z")

</div>

No, local machine only have 16GB RAM, and there’re 25 timesteps in the file.

---

<div class="post-metadata">

### Author: ![ZiweiWu](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/z/ea666f/32.png) [@ZiweiWu](https://discourse.paraview.org/u/ZiweiWu)
#### Post date: [November 8, 2018, 3:24am UTC](https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860/5 "2018-11-08T03:24:41Z")

</div>

Yeah, this could be a solution, I’ll try it out. And yes, I’m using the PVGeo-CMAQ reader, the file contains 25 timesteps (looks like the file I sent you last time but has bigger scale). Do I have to update script in both PVGeo and PVGeo-HDF5 or only update PVGeo-HDF5?  
Thanks!

---

<div class="post-metadata">

### Author: ![banesullivan](https://discourse.paraview.org/user_avatar/discourse.paraview.org/banesullivan/32/11714_2.png) [@banesullivan](https://discourse.paraview.org/u/banesullivan)
#### Post date: [November 8, 2018, 6:57pm UTC](https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860/6 "2018-11-08T18:57:15Z")

</div>

You will only need to update the code for the PVGeo-CMAQ reader in [PVGeo-HDF](https://github.com/OpenGeoVis/PVGeo-HDF5).

To outline the changes needed, instead of reading the data up front (this is what the PVGeo-CMAQ reader is currently doing), you’ll update it to get the requested time step and pass that timestep to the reading function to only read the requested part of the file.

So change lines [190-193](https://github.com/OpenGeoVis/PVGeo-HDF5/blob/master/pvgeohdf/netcdf.py#L191) to pass the requested timestep on every `RequestData` call. Then change [`_ReadUpFront`](https://github.com/OpenGeoVis/PVGeo-HDF5/blob/master/pvgeohdf/netcdf.py#L126) to take the timestep index and restructure that function to only grab the needed data for each time step.

This might be a bit clunky, but it should allow you to visualize the 50GB data file on your local machine.

---

<div class="post-metadata">

### Author: ![ZiweiWu](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/z/ea666f/32.png) [@ZiweiWu](https://discourse.paraview.org/u/ZiweiWu)
#### Post date: [November 11, 2018, 1:33am UTC](https://discourse.paraview.org/t/import-giant-50gb-data-file-into-paraview/860/7 "2018-11-11T01:33:53Z")

</div>

Got it, I’ll try it out later. Thanks Bane!
