Infrastructure
In this section:
This section describes the equipment that makes up GridUnesp and the structure that supports its operation.
Data Center
GridUnesp is located in the Data Center of the Center for Scientific Computing of Unesp (NCC/Unesp), on the Barra Funda campus, in São Paulo. This data center hosts not only GridUnesp but also other projects of strategic importance to Unesp and the scientific community.
Projects in the Data Center
SPRACE (São Paulo Research and Analysis Center)
SPRACE operates a Tier-2 of the Worldwide LHC Computing Grid (WLCG)
Processing and analysis of data from the CMS experiment at CERN
GridUnesp was born as a spin-off of SPRACE
ANSPGrid CA (Academic Network at São Paulo Grid Certification Authority)
The ANSPGrid CA issues X509 grid certificates for researchers since 2013
Based on public-key/asymmetric cryptography
Serves research institutes throughout Brazil since 2018
Modern Code & Intel CoE for Machine Learning
Cluster based on the Intel Many Integrated Core (MIC) architecture
Training in OpenMP, MPI, and parallel programming
Mirror of the Rectory
Backup of Unesp’s administrative systems
Ensures continuous availability of services
Unesp Institutional Repository
The Unesp Institutional Repository is the Digital Library
Storage and preservation of Unesp’s scientific output
Open access to documents, data, and management plans
unespNET
unespNET is connected to the Rectory and to Brazil’s NAP
Connection over optical channels at 100 Gbps
Receives links from interior units (Presidente Prudente, Bauru, and Botucatu)
Interconnects 34 university units in 24 cities
Approximately 60,000 users
The NCC Data Center as a strategic point of presence
TnCentral Database and Genomic Research
Deployed in 2021 in the NCC data center
-
biological data on prokaryotic transposable elements related to
bacterial resistance to antibiotics
disease outbreaks in economically important crops and livestock
public data and “genome browsers” for all plant genomes developed by the
Laboratory of Genomics and Bioinformatics of Plants and Microorganisms
Laboratory of Plant Systematics and Evolution at UNESP-FCAV
International collaboration with several institutions
Unesp
Protein Information Resource
Multidrug-Resistant Organism Repository and Surveillance Network
Walter Reed Army Institute of Research
Georgetown University Medical Center
University of Delaware
Team
The GridUnesp team, integrated with the other NCC projects, brings together specialists in several areas of computing:
Networking and connectivity
Hardware and infrastructure
Software development
Parallel programming and optimization
High-performance computing
Data analysis
Information security
Systems auditing
Team responsibilities:
24/7 operation of the Data Center
Acquisition and maintenance of equipment
Management of internet connections and electrical power
Cooling and security systems
Support for the scientific community
Administration and management of resources
In addition to the technical team, the NCC has qualified administrative staff for budget and bureaucratic management.
The GridUnesp Cluster
Definitions
Cluster: A set of computers working simultaneously on large, repetitive tasks
Grid: A set of clusters interconnected in a coordinated manner
GridUnesp began operations as a grid, but currently operates as a cluster due to the instability of the KyaTera network that connected the campuses.
Operation
In operation since 2009
Runs 24 hours a day, 7 days a week, 365 days a year
Sporadic downtimes only for installing new equipment and preventive maintenance
Some original equipment still in operation (more than 15 years of use)
Node Configuration
Worker Nodes (Processing Nodes)
GridUnesp has 57 processing nodes:
56 CPU nodes (node001 to node056)
1 GPU node (gpunode001)
Node Type |
Total Capacity |
Reserved for OS |
Available for Jobs |
Regular Nodes ( |
28 Cores / 56 CPUs |
2 Cores / 4 CPUs |
26 Cores / 52 CPUs |
GPU Node ( |
48 Cores / 96 CPUs |
4 Cores / 8 CPUs |
44 Cores / 88 CPUs |
Specifications of the CPU nodes (node001 - node056):
CPU:
Manufacturer: Intel
Model: Intel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz
Architecture: x86_64 (64-bit)
Physical cores: 14 per socket
Sockets: 2 (28 cores per node in total)
Threads: 28 (Hyper-Threading enabled)
Frequency: 2.40 GHz (base)
L3 cache: 35 MB
Memory:
RAM: 128 GB per node (4 GB/core)
Features:
- Support for AVX2, FMA
- Virtualization (VMX)
- Intel Turbo Boost
- Intel Hyper-Threading
Specifications of the GPU node (gpunode001):
CPU:
Manufacturer: AMD
Model: AMD EPYC 9224 24-Core Processor
Architecture: x86_64 (64-bit)
Physical cores: 24 per socket
Sockets: 2 (48 cores per node in total)
Threads: 96 (SMT enabled)
Frequency: 3.7 GHz (boost)
L3 cache: 1024 KB per core
GPU:
Manufacturer: NVIDIA
Model: NVIDIA L40S
Quantity: 4 GPUs per node
Architecture: Ada Lovelace
Memory per GPU: 48 GB GDDR6
Total GPU memory: 192 GB
CUDA Capability: 8.9
CUDA Cores: 18,176 per GPU
Tensor Cores: 568 per GPU (4th generation)
RT Cores: 142 per GPU (3rd generation)
Memory:
RAM: 1.5 TB
Features:
- Support for AVX-512
- AMD-V (virtualization)
- PCIe 4.0/5.0
- Full support for CUDA, cuDNN, TensorRT
- Optimized for Deep Learning, AI, and rendering
Note
For more information on using GPUs, see the GPU Usage section.
Specifications Summary
Feature |
Previously |
After upgrade |
Processing (TF) |
23 |
77 |
Storage (TB) |
136 |
288 |
Cores |
2,048 |
3,232 |
Upgrade details:
56 processing nodes: 1,568 cores, 4 GB/core
1 GPU node: AMD EPYC 9224 processor
4 management servers: 12 cores/server, 40 GbE network
Storage: 288 TB raw (216 TB effective) with Lustre
Internal network: 40 GbE between nodes, 2 switches with 32 ports
Warning
Storage Policy:
There are no disk quotas per user or project
There is no automatic backup system
Users are responsible for their own backups
In case of a catastrophic failure, data will be lost with no possibility of recovery.
Network
Currently, GridUnesp uses a 40 GB/s Ethernet network for:
Communication between processing nodes
Communication with storage
Note
Previously, the cluster used an InfiniBand network, considered faster but with maintenance costs that were unsustainable in the long term. The migration to Ethernet occurred with the 2017/2018 upgrade.
Energy Consumption
The 2017/2018 upgrade provided a significant reduction in energy consumption:
Before: 90 kW/h
After: 15.8 kW/h
Electrical infrastructure:
2 UPS units to support power outages
Diesel generator with autonomy for hours of operation
Scheduled sequential shutdown in case of prolonged failure
Cooling System
The NCC/Unesp Data Center has a high-capacity cooling system, designed to maintain the ideal temperature and humidity conditions for the continuous and safe operation of the equipment.
System components:
3 (three) precision air-conditioning units, with capacity suited to the thermal load of the environment
12 (twelve) large fans located on the exterior of the building, responsible for dissipating the heat generated by the equipment
Cooling system
Containment of cold air in the aisle where the most intensive production equipment is concentrated (cold aisle), which provides:
Increased thermal efficiency of the cooling system
Reduced energy consumption
Better temperature control in high-density racks
Future Acquisitions
Incorporation Policy
Maintaining an infrastructure like GridUnesp’s requires:
qualified staff
updating of equipment due to natural degradation and inevitable obsolescence
The Center for Scientific Computing of Unesp has always sought funding opportunities from research agencies to acquire new equipment to provide the scientific community with greater storage capacity, processing speed, and data transfer. With these actions, GridUnesp seeks to ensure that computing resources continue to drive research at the university, opening new frontiers for scientific investigation at Unesp.
Attention
Accordingly, GridUnesp has implemented a new policy for incorporating computing resources:
Projects awarded new High-Performance Computing (HPC) equipment can request incorporation into the Data Center
Priority of use for 3 years for the incorporated resources
After this period, equipment is incorporated into the shared structure
Benefits:
Optimization of financial resources (the Data Center cost is shared)
Continuous renewal of the infrastructure
Access to the latest technologies
Unsupported Services
To better plan your work, it is important to be aware of the services that are not available:
Data Transfer:
Globus/GridFTP: Not supported
Alternatives:
scp,rsync,sftp, WinSCP (see Transferring Files)
Interactive Environments:
Jupyter Notebooks: Not supported
ParaView: Not supported in interactive mode
Alternative: Run analyses and visualizations on your local machine
Authentication:
Two-Factor Authentication (2FA): Not available
Access is done exclusively via SSH password
Storage:
Automatic backup: Not available
Disk quotas: Not defined (use responsibly)
Important
Data backup is the user’s responsibility.
Since there is no backup system, it is essential that you:
Make regular backups of important data
Keep copies in at least two different locations
Do not use GridUnesp as the only storage location for critical data
See also
Storage Information (Detailed) - Details about storage
Best Practices - Usage recommendations
Accessing the Cluster - How to access the cluster