Infrastructure

This section describes the equipment that makes up GridUnesp and the structure that supports its operation.

Data Center

NCC/Unesp Data Center

GridUnesp is located in the Data Center of the Center for Scientific Computing of Unesp (NCC/Unesp), on the Barra Funda campus, in São Paulo. This data center hosts not only GridUnesp but also other projects of strategic importance to Unesp and the scientific community.

Projects in the Data Center

SPRACE (São Paulo Research and Analysis Center)

  • SPRACE operates a Tier-2 of the Worldwide LHC Computing Grid (WLCG)

  • Processing and analysis of data from the CMS experiment at CERN

  • GridUnesp was born as a spin-off of SPRACE

ANSPGrid CA (Academic Network at São Paulo Grid Certification Authority)

  • The ANSPGrid CA issues X509 grid certificates for researchers since 2013

  • Based on public-key/asymmetric cryptography

  • Serves research institutes throughout Brazil since 2018

Modern Code & Intel CoE for Machine Learning

  • Cluster based on the Intel Many Integrated Core (MIC) architecture

  • Training in OpenMP, MPI, and parallel programming

Mirror of the Rectory

  • Backup of Unesp’s administrative systems

  • Ensures continuous availability of services

Unesp Institutional Repository

  • The Unesp Institutional Repository is the Digital Library

  • Storage and preservation of Unesp’s scientific output

  • Open access to documents, data, and management plans

unespNET

  • unespNET is connected to the Rectory and to Brazil’s NAP

  • Connection over optical channels at 100 Gbps

  • Receives links from interior units (Presidente Prudente, Bauru, and Botucatu)

  • Interconnects 34 university units in 24 cities

  • Approximately 60,000 users

  • The NCC Data Center as a strategic point of presence

TnCentral Database and Genomic Research

  • Deployed in 2021 in the NCC data center

  • Database

    • biological data on prokaryotic transposable elements related to

      • bacterial resistance to antibiotics

      • disease outbreaks in economically important crops and livestock

    • public data and “genome browsers” for all plant genomes developed by the

      • Laboratory of Genomics and Bioinformatics of Plants and Microorganisms

      • Laboratory of Plant Systematics and Evolution at UNESP-FCAV

  • International collaboration with several institutions

    • Unesp

    • Protein Information Resource

    • Multidrug-Resistant Organism Repository and Surveillance Network

    • Walter Reed Army Institute of Research

    • Georgetown University Medical Center

    • University of Delaware

Team

The GridUnesp team, integrated with the other NCC projects, brings together specialists in several areas of computing:

  • Networking and connectivity

  • Hardware and infrastructure

  • Software development

  • Parallel programming and optimization

  • High-performance computing

  • Data analysis

  • Information security

  • Systems auditing

Team responsibilities:

  • 24/7 operation of the Data Center

  • Acquisition and maintenance of equipment

  • Management of internet connections and electrical power

  • Cooling and security systems

  • Support for the scientific community

  • Administration and management of resources

In addition to the technical team, the NCC has qualified administrative staff for budget and bureaucratic management.

The GridUnesp Cluster

Definitions

  • Cluster: A set of computers working simultaneously on large, repetitive tasks

  • Grid: A set of clusters interconnected in a coordinated manner

GridUnesp began operations as a grid, but currently operates as a cluster due to the instability of the KyaTera network that connected the campuses.

Operation

  • In operation since 2009

  • Runs 24 hours a day, 7 days a week, 365 days a year

  • Sporadic downtimes only for installing new equipment and preventive maintenance

  • Some original equipment still in operation (more than 15 years of use)

Node Configuration

Worker Nodes (Processing Nodes)

GridUnesp has 57 processing nodes:

  • 56 CPU nodes (node001 to node056)

  • 1 GPU node (gpunode001)

Cluster Resource Allocation (Core Specialization)

Node Type

Total Capacity

Reserved for OS

Available for Jobs

Regular Nodes (node[001-056])

28 Cores / 56 CPUs

2 Cores / 4 CPUs

26 Cores / 52 CPUs

GPU Node (gpunode001)

48 Cores / 96 CPUs

4 Cores / 8 CPUs

44 Cores / 88 CPUs

Specifications of the CPU nodes (node001 - node056):

CPU:
  Manufacturer: Intel
  Model: Intel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz
  Architecture: x86_64 (64-bit)
  Physical cores: 14 per socket
  Sockets: 2 (28 cores per node in total)
  Threads: 28 (Hyper-Threading enabled)
  Frequency: 2.40 GHz (base)
  L3 cache: 35 MB

Memory:
  RAM: 128 GB per node (4 GB/core)

Features:
  - Support for AVX2, FMA
  - Virtualization (VMX)
  - Intel Turbo Boost
  - Intel Hyper-Threading

Specifications of the GPU node (gpunode001):

CPU:
  Manufacturer: AMD
  Model: AMD EPYC 9224 24-Core Processor
  Architecture: x86_64 (64-bit)
  Physical cores: 24 per socket
  Sockets: 2 (48 cores per node in total)
  Threads: 96 (SMT enabled)
  Frequency: 3.7 GHz (boost)
  L3 cache: 1024 KB per core

GPU:
  Manufacturer: NVIDIA
  Model: NVIDIA L40S
  Quantity: 4 GPUs per node
  Architecture: Ada Lovelace
  Memory per GPU: 48 GB GDDR6
  Total GPU memory: 192 GB
  CUDA Capability: 8.9
  CUDA Cores: 18,176 per GPU
  Tensor Cores: 568 per GPU (4th generation)
  RT Cores: 142 per GPU (3rd generation)

Memory:
  RAM: 1.5 TB

Features:
  - Support for AVX-512
  - AMD-V (virtualization)
  - PCIe 4.0/5.0
  - Full support for CUDA, cuDNN, TensorRT
  - Optimized for Deep Learning, AI, and rendering

Note

For more information on using GPUs, see the GPU Usage section.

Specifications Summary

Feature

Previously

After upgrade

Processing (TF)

23

77

Storage (TB)

136

288

Cores

2,048

3,232

Upgrade details:

  • 56 processing nodes: 1,568 cores, 4 GB/core

  • 1 GPU node: AMD EPYC 9224 processor

  • 4 management servers: 12 cores/server, 40 GbE network

  • Storage: 288 TB raw (216 TB effective) with Lustre

  • Internal network: 40 GbE between nodes, 2 switches with 32 ports

Warning

Storage Policy:

  • There are no disk quotas per user or project

  • There is no automatic backup system

  • Users are responsible for their own backups

In case of a catastrophic failure, data will be lost with no possibility of recovery.

Network

Currently, GridUnesp uses a 40 GB/s Ethernet network for:

  • Communication between processing nodes

  • Communication with storage

Note

Previously, the cluster used an InfiniBand network, considered faster but with maintenance costs that were unsustainable in the long term. The migration to Ethernet occurred with the 2017/2018 upgrade.

Energy Consumption

The 2017/2018 upgrade provided a significant reduction in energy consumption:

  • Before: 90 kW/h

  • After: 15.8 kW/h

Electrical infrastructure:

  • 2 UPS units to support power outages

  • Diesel generator with autonomy for hours of operation

  • Scheduled sequential shutdown in case of prolonged failure

Cooling System

The NCC/Unesp Data Center has a high-capacity cooling system, designed to maintain the ideal temperature and humidity conditions for the continuous and safe operation of the equipment.

System components:

  • 3 (three) precision air-conditioning units, with capacity suited to the thermal load of the environment

  • 12 (twelve) large fans located on the exterior of the building, responsible for dissipating the heat generated by the equipment

    Cooling system
    _images/ventoinha1.jpg _images/ventoinha2.jpg
  • Containment of cold air in the aisle where the most intensive production equipment is concentrated (cold aisle), which provides:

    • Increased thermal efficiency of the cooling system

    • Reduced energy consumption

    • Better temperature control in high-density racks

Future Acquisitions

Incorporation Policy

Maintaining an infrastructure like GridUnesp’s requires:

  • qualified staff

  • updating of equipment due to natural degradation and inevitable obsolescence

The Center for Scientific Computing of Unesp has always sought funding opportunities from research agencies to acquire new equipment to provide the scientific community with greater storage capacity, processing speed, and data transfer. With these actions, GridUnesp seeks to ensure that computing resources continue to drive research at the university, opening new frontiers for scientific investigation at Unesp.

Attention

Accordingly, GridUnesp has implemented a new policy for incorporating computing resources:

  1. Projects awarded new High-Performance Computing (HPC) equipment can request incorporation into the Data Center

  2. Priority of use for 3 years for the incorporated resources

  3. After this period, equipment is incorporated into the shared structure

Benefits:

  • Optimization of financial resources (the Data Center cost is shared)

  • Continuous renewal of the infrastructure

  • Access to the latest technologies

Unsupported Services

To better plan your work, it is important to be aware of the services that are not available:

Data Transfer:

  • Globus/GridFTP: Not supported

  • Alternatives: scp, rsync, sftp, WinSCP (see Transferring Files)

Interactive Environments:

  • Jupyter Notebooks: Not supported

  • ParaView: Not supported in interactive mode

  • Alternative: Run analyses and visualizations on your local machine

Authentication:

  • Two-Factor Authentication (2FA): Not available

  • Access is done exclusively via SSH password

Storage:

  • Automatic backup: Not available

  • Disk quotas: Not defined (use responsibly)

Important

Data backup is the user’s responsibility.

Since there is no backup system, it is essential that you:

  1. Make regular backups of important data

  2. Keep copies in at least two different locations

  3. Do not use GridUnesp as the only storage location for critical data

See also