# pvbatch with EGL device on HPC

**URL:** https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262
**Category:** ParaView Support
**Tags:** opengl
**Created:** [March 23, 2022, 10:17am UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262 "2022-03-23T10:17:03Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![jirikolar](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/j/97f17d/32.png) [@jirikolar](https://discourse.paraview.org/u/jirikolar)
#### Post date: [March 23, 2022, 10:17am UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262/1 "2022-03-23T10:17:03Z")

</div>

Hi,  
I compiled paraview in headless mode with EGL support. However, I can use it only if PBS scheduler assigns me the very first GPU in the node (index 0). If it assigns me any other card, despite it is visible within the job as the device with index 0, the EGL device can not be initialized.

Do you have any idea what might be wrong?

Additionally, I have found no way to specify the EGL device index manually, I always tried variable $CUDA\_VISIBLE\_DEVICES or the index according to nvidia-smi:

> pvbatch --egl-device-index=$CUDA\_VISIBLE\_DEVICES test.py  
> pvbatch --egl-device-index=2 test.py  
> Error: unknown option --egl-device-index=…

> pvbatch test.py -display=$CUDA\_VISIBLE\_DEVICES  
> Not working

> pvbatch --displays=$CUDA\_VISIBLE\_DEVICES test.py  
> Error No script specified. Please specify a batch script

If I am assigned the very first GPU, then, without problems, I can just run:

> pvbatch test.py

Paraview was compiled with success using:

> cmake -GNinja -DCMAKE\_BUILD\_TYPE=Release -DPARAVIEW\_USE\_PYTHON=ON -DPARAVIEW\_USE\_MPI=ON -DVTK\_SMP\_IMPLEMENTATION\_TYPE=TBB -DVTK\_OPENGL\_HAS\_EGL=ON -DVTK\_OPENGL\_HAS\_OSMESA=OFF -DPARAVIEW\_USE\_QT=OFF -DVTK\_USE\_X=OFF -DEGL\_opengl\_LIBRARY:FILEPATH=/usr/lib/x86\_64-linux-gnu/libOpenGL.so.0 -DEGL\_INCLUDE\_DIR=/usr/include/EGL/ -DEGL\_LIBRARY=/usr/lib/x86\_64-linux-gnu/nvidia/current/libEGL\_nvidia.so.0 -DTBB\_ROOT=$TBBROOT -DPARAVIEW\_ENABLE\_FFMPEG=ON …/…/paraview  
> ninja -j4

Thank you for your help.  
Jiri

---

<div class="post-metadata">

### Author: ![utkarsh.ayachit](https://discourse.paraview.org/user_avatar/discourse.paraview.org/utkarsh.ayachit/32/39_2.png) [@utkarsh.ayachit](https://discourse.paraview.org/u/utkarsh.ayachit)
#### Post date: [March 23, 2022, 12:21pm UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262/2 "2022-03-23T12:21:54Z")

</div>

Try `pvbatch --displays=$CUDA_VISIBLE_DEVICES -- test.py`.

---

<div class="post-metadata">

### Author: ![jirikolar](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/j/97f17d/32.png) [@jirikolar](https://discourse.paraview.org/u/jirikolar)
#### Post date: [March 23, 2022, 12:48pm UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262/3 "2022-03-23T12:48:24Z")

</div>

Now it can see the script, but the resulting error with EGL is still the same:

> ( 1.100s) [pvbatch] vtkEGLRenderWindow.cxx:353 WARN| vtkEGLRenderWindow (0x556bb0976040): EGL device index: 0 could not be initialized. Trying other devices…  
> ( 1.105s) [pvbatch] vtkEGLRenderWindow.cxx:386 WARN| vtkEGLRenderWindow (0x555f5634a070): Setting an EGL display to device index: -1 require EGL\_EXT\_device\_base EGL\_EXT\_platform\_device EGL\_EXT\_platform\_base extensions  
> ( 1.105s) [pvbatch] vtkEGLRenderWindow.cxx:388 WARN| vtkEGLRenderWindow (0x555f5634a070): Attempting to use EGL\_DEFAULT\_DISPLAY…  
> ( 1.105s) [pvbatch] vtkEGLRenderWindow.cxx:393 ERR| vtkEGLRenderWindow (0x555f5634a070): Could not initialize a device. Exiting…  
> ( 1.105s) [pvbatch]vtkOpenGLRenderWindow.c:511 ERR| vtkEGLRenderWindow (0x555f5634a070): GLEW could not be initialized: Missing GL version

Regardless of what I supply to the display, it attempts to use device index 0. This should be correct, as PBS gives me resources in a separate namespace, in nvidia-smi output, there is only one GPU visible, with index 0. But when I ssh directly to the cluster and use nvidia-smi, there are 8 different GPUs and for example, what I see from the PBS job as index 0, is in fact GPU with index 7 within whole machine.

I assume that problem is that paraview tries to use EGL by index and not by persistent address, e.g. Bus-ID, or the ID of the GPU which is in the $CUDA\_VISIBLE\_DEVICES variable.

---

<div class="post-metadata">

### Author: ![jirikolar](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/j/97f17d/32.png) [@jirikolar](https://discourse.paraview.org/u/jirikolar)
#### Post date: [March 23, 2022, 12:50pm UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262/4 "2022-03-23T12:50:51Z")

</div>

According to bus-id (00000000:C2:00.0) I was assigned GPU with id number 7.

nvidia-smi (in PBS job namespace)

> ±----------------------------------------------------------------------------+  
> | NVIDIA-SMI 460.73.01 Driver Version: 460.73.01 CUDA Version: 11.2 |  
> |-------------------------------±---------------------±---------------------+  
> | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |  
> | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |  
> | | | MIG M. |  
> |===============================|  
> | 0 RTX A4000 On | 00000000:C2:00.0 Off | Off |  
> | 41% 26C P8 15W / 140W | 1MiB / 16117MiB | 0% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+

nvidia-smi for the whole node

> ±----------------------------------------------------------------------------+  
> | NVIDIA-SMI 460.73.01 Driver Version: 460.73.01 CUDA Version: 11.2 |  
> |-------------------------------±---------------------±---------------------+  
> | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |  
> | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |  
> | | | MIG M. |  
> |===============================|  
> | 0 RTX A4000 On | 00000000:01:00.0 Off | Off |  
> | 41% 42C P2 42W / 140W | 2260MiB / 16117MiB | 13% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+  
> | 1 RTX A4000 On | 00000000:21:00.0 Off | Off |  
> | 41% 25C P8 14W / 140W | 1MiB / 16117MiB | 0% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+  
> | 2 RTX A4000 On | 00000000:22:00.0 Off | Off |  
> | 41% 27C P8 15W / 140W | 1MiB / 16117MiB | 0% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+  
> | 3 RTX A4000 On | 00000000:41:00.0 Off | Off |  
> | 47% 65C P2 86W / 140W | 14882MiB / 16117MiB | 23% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+  
> | 4 RTX A4000 On | 00000000:81:00.0 Off | Off |  
> | 41% 26C P8 13W / 140W | 1MiB / 16117MiB | 0% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+  
> | 5 RTX A4000 On | 00000000:A1:00.0 Off | Off |  
> | 41% 42C P2 45W / 140W | 2258MiB / 16117MiB | 13% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+  
> | 6 RTX A4000 On | 00000000:C1:00.0 Off | Off |  
> | 41% 28C P8 13W / 140W | 1MiB / 16117MiB | 0% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+  
> | 7 RTX A4000 On | 00000000:C2:00.0 Off | Off |  
> | 41% 25C P8 15W / 140W | 1MiB / 16117MiB | 0% Default |  
> | | | N/A |  
> ±------------------------------±---------------------±---------------------+

---

<div class="post-metadata">

### Author: ![utkarsh.ayachit](https://discourse.paraview.org/user_avatar/discourse.paraview.org/utkarsh.ayachit/32/39_2.png) [@utkarsh.ayachit](https://discourse.paraview.org/u/utkarsh.ayachit)
#### Post date: [March 23, 2022, 6:29pm UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262/5 "2022-03-23T18:29:30Z")

</div>

> [@jirikolar](#):
>
> I assume that problem is that paraview tries to use EGL by index and not by persistent address, e.g. Bus-ID, or the ID of the GPU which is in the $CUDA\_VISIBLE\_DEVICES variable.

@danlipsa, do you know if this is indeed the case?

---

<div class="post-metadata">

### Author: ![danlipsa](https://discourse.paraview.org/user_avatar/discourse.paraview.org/danlipsa/32/1563_2.png) [@danlipsa](https://discourse.paraview.org/u/danlipsa)
#### Post date: [March 23, 2022, 6:39pm UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262/6 "2022-03-23T18:39:54Z")

</div>

Indeed that is the case. We get a list of available devices from EGL and we use the one with the specified index.

---

<div class="post-metadata">

### Author: ![jirikolar](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/j/97f17d/32.png) [@jirikolar](https://discourse.paraview.org/u/jirikolar)
#### Post date: [March 23, 2022, 7:31pm UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262/7 "2022-03-23T19:31:29Z")

</div>

Well, I really don’t know what is happening, this is just my guess. From what I tested, on the same machine, EGL could be initialized when I was assigned the first GPU card and it could not if I was assigned any other one. This I have verified on three different machines, same behavior.

---

<div class="post-metadata">

### Author: ![jirikolar](https://discourse.paraview.org/letter_avatar_proxy/v4/letter/j/97f17d/32.png) [@jirikolar](https://discourse.paraview.org/u/jirikolar)
#### Post date: [March 23, 2022, 8:51pm UTC](https://discourse.paraview.org/t/pvbatch-with-egl-device-on-hpc/9262/8 "2022-03-23T20:51:18Z")

</div>

So, after consulting our HPC admin, it seems that the problem is a bug in the GPU driver.

Problematic machines:

> Driver Version: 460.73.01 CUDA Version: 11.2

Newer HPC machines with no issues:

> Driver Version: 470.103.01 CUDA Version: 11.4
