|
Book Cephas, Our 160-Core, ~7.0 TFLOPS ADCIRC Live Cluster
Our Services
Access to our managed ADCIRC cluster is not free, but it is very cost-effective. It is not running in the cloud; it is a physical cluster that we operate and maintain in Houston, Texas, USA.
We call our dependable, bare-metal computing cluster, Cephas.
We are continuing to refine the best ways to provide access and support. For now, interested users should
email us
to discuss their requirements.
We can reserve blocks of days or weeks during which a client may use the cluster exclusively. The exception is an operational emergency, such as a hurricane, when we may need to reclaim the cluster for forecasting or response work.
Exclusive access to Cephas starts at $100 per day, with a minimum reservation of three days. Final pricing depends on the reservation length, scheduling priority, storage and data-transfer requirements, and the amount of technical or engineering support requested.
| Service |
Pricing |
| Self-managed full-cluster access for 7 or more days |
Starting at $100/day |
| Self-managed full-cluster access for 3–6 days |
$125–$150/day |
| Managed ADCIRC runs |
$175–$250/day |
| Engineering, model setup, or troubleshooting support |
$75–$125/hour |
| Urgent or priority reservations |
Custom quote |
Self-managed pricing covers reserved computing access and normal cluster operation. Model preparation, ADCIRC configuration, troubleshooting, data processing, engineering assistance, and other hands-on services are quoted separately.
We work best with friendly, collaborative clients who share our goal of getting their ADCIRC simulations completed successfully.
Email us
to book Cephas,
with full-cluster access starting at $100 per day.
About Cephas
Cephas is our four-node, 160-physical-core ADCIRC computing cluster, connected by a private 10 Gigabit Ethernet fabric for inter-node MPI communication and NFS shared-storage traffic. Rather than nickel-and-dime users with complicated CPU-hour pricing, we prefer to learn about each client’s requirements and provide a customized quote.
We are open to renting part or all of the cluster for fixed blocks of time, depending on scheduling and priority. We use Cephas for our own ADCIRC work as well, but our rental rates can be substantially lower than comparable commercial cloud-computing costs.
Our goal is the same as yours: to get your ADCIRC runs completed. Tell us about your mesh, simulation duration, forcing, number of scenarios, output requirements, and desired completion date, and we will help determine the most practical way to run the work.
You can also jump to
the comparison below
to see approximately what would have been required to provide comparable theoretical computing capacity in 2006, shortly after Hurricane Katrina.
Email us at
help@support.adcirc.live
to request information about renting Cephas and receive a free quote!
Sign up for our
announcement email list
to receive updates about cluster availability, capabilities, and services.
Cephas: Our ~7.0 TFLOPS ADCIRC Cluster Specifications
| Component |
Details |
| Compute Nodes |
4 × Dell Precision T7810 workstations |
| Physical CPU Cores |
40 physical cores per node
160 physical cores cluster-wide
|
| Processors |
Each node: 2 × Intel Xeon E5-2698 v4 processors
Each processor: 20 physical cores, 2.2 GHz base frequency, up to 3.6 GHz turbo
Total per node: 2 processors and 40 physical cores
Complete cluster: 8 processors and 160 physical cores
|
| Memory |
compute01: 32 GB DDR4 RAM
compute02: 32 GB DDR4 RAM
compute03: 32 GB DDR4 RAM
compute04: 32 GB DDR4 RAM
Total installed physical memory: 128 GB
Slurm configured RealMemory: 31,000 MB per node, or 124,000 MB cluster-wide
|
| Private Compute Network |
Dedicated switched 10 Gigabit Ethernet fabric connecting all four compute nodes
Used for inter-node Intel MPI communication and NFS shared-storage traffic
Keeps high-volume compute and file traffic separate from the normal management network
|
| Shared Storage |
16 TB shared storage
NFS-based shared user, software, model-input, and simulation-output directories
NFS traffic is carried across the private 10 GbE fabric
|
| Job Scheduler |
Slurm workload manager
Supports single-node and multi-node ADCIRC jobs using up to 160 tasks
Full-cluster runs may use 159 computational MPI ranks plus one dedicated ADCIRC output-writer task
|
| Software |
Ubuntu Linux
Parallel ADCIRC and ADCIRC+SWAN builds
Intel oneAPI HPC Compiler Suite
Intel MPI
NetCDF and HDF5 scientific-data libraries
Docker services available on compute01
|
| Estimated Peak Performance |
Estimated double-precision theoretical peak: approximately 0.88 TFLOPS per processor.
8 processors × approximately 0.88 TFLOPS = approximately 7.0 TFLOPS FP64 cluster-wide.
Estimated single-precision theoretical peak is approximately twice that amount, or approximately 14.1 TFLOPS FP32.
A current post-upgrade EC2001 benchmark using a 254,565-vertex mesh, 159 computational MPI ranks, and one dedicated writer completed 1,000 one-second time steps in approximately 8.06 seconds. That is approximately 124 times faster than real time, corresponding to about 5.8 hours of solver time for a 30-day simulation if that rate is sustained.
These figures are theoretical estimates and representative benchmark results, not guarantees. Actual ADCIRC performance depends on mesh size, communication requirements, output frequency, storage activity, forcing, compiler settings, and how efficiently the workload scales across the four nodes. We did not retain a directly comparable pre-upgrade benchmark, so we do not claim a precise percentage speedup attributable only to the network.
|
ADCIRC Capacity Rule of Thumb:
As a rough planning estimate, ADCIRC may use approximately 5,000 mesh nodes per physical CPU core.
• Cephas capacity at 160 physical cores: approximately 800,000 mesh nodes.
This is not a hard limit. Practical capacity and performance depend on the ADCIRC configuration, mesh characteristics, meteorological forcing, wave coupling, output requirements, available memory, and desired turnaround time.
Recent Improvements and Planned Upgrades
Cephas is an actively maintained system. We have expanded the cluster to four compute nodes, 160 physical CPU cores, and 128 GB of installed memory, while also completing a major network upgrade for both MPI communication and NFS shared-storage traffic.
-
128 GB of balanced cluster memory:
All four compute nodes now contain 32 GB of DDR4 ECC memory, providing 128 GB of installed physical RAM cluster-wide. Memory is distributed evenly across the four nodes, and Slurm is configured with 31,000 MB of RealMemory per node, providing consistent scheduling capacity across compute01 through compute04.
-
Operational private 10 Gigabit Ethernet fabric:
All four compute nodes are now connected through a dedicated switched 10 GbE network. Intel MPI inter-node communication and NFS shared-storage traffic use this private fabric instead of competing on the former Gigabit Ethernet path.
-
More bandwidth for multi-node ADCIRC:
The upgraded links provide ten times the nominal line rate of Gigabit Ethernet: 10 Gb/s, or 1.25 GB/s theoretical, per link instead of 1 Gb/s, or 125 MB/s theoretical. After protocol overhead, the network-side ceiling for a large sequential transfer rises from roughly 110 MB/s on 1 GbE to approximately 1 GB/s on 10 GbE, provided that the source and destination storage can sustain that rate.
-
Reduced MPI communication delays:
ADCIRC exchanges subdomain-boundary information throughout a parallel run. The faster private fabric reduces the amount of time ranks spend waiting on inter-node communication and improves the ability of a 160-task job to make useful use of all four nodes. The benefit is greatest for meshes and processor counts where communication had begun to limit scaling.
-
Improved NFS startup and output behavior:
Model inputs, executables, shared software, restart files, and simulation output can move across the same 10 GbE fabric. This reduces network contention during model startup, NetCDF or HDF5 output, checkpointing, post-processing, and repeated ensemble workflows. Actual NFS throughput can still be limited by the current hard-disk storage, file size, access pattern, and the number of simultaneous readers or writers.
-
Current full-cluster runtime:
In a post-upgrade EC2001 test with 254,565 mesh vertices, 159 computational MPI ranks, and one dedicated writer, Cephas completed 1,000 one-second time steps in approximately 8.06 seconds. This is about 124 times faster than real time and projects to approximately 5.8 hours of solver time for a 30-day simulation if the same rate is maintained. Startup, forcing, output frequency, and post-processing can change total wall-clock time.
-
Careful performance claims:
We did not retain a directly comparable benchmark from the former 1 GbE network, so we do not claim that the entire observed runtime is a specific percentage faster solely because of 10 GbE. The upgrade does remove the old 1 Gb/s network ceiling and provides substantially more headroom for MPI scaling and shared-storage I/O.
-
Improved Slurm scheduling topology:
We are continuing to refine the Slurm configuration so that cluster-management, login, scheduling, storage, and compute responsibilities are clearly separated. Job placement can also account for node memory, network communication, shared-storage access, and the location of ADCIRC output-writing tasks.
-
Faster shared storage:
Further down the road, we plan to supplement or replace portions of the current spinning-hard-drive storage with solid-state storage. High-speed NVMe storage is the preferred long-term performance option because it can take better advantage of the new 10 GbE network while providing substantially higher throughput and lower latency than conventional hard disks.
-
Tiered storage:
Large-capacity hard disks may continue to be used for archival and less performance-sensitive data, while faster SSD or NVMe storage could be used for active model inputs, temporary files, checkpoints, and simulation output.
The network upgrade is intended to improve real ADCIRC turnaround time rather than merely increase a theoretical specification. It is especially valuable for full-cluster simulations, output-heavy jobs, repeated ensembles, shared-file workflows, and cases in which MPI communication or NFS traffic would otherwise compete for a 1 GbE link.
If you need a cost-effective environment in which to scale your ADCIRC capacity quickly,
click here to request a free quote
!
ADCIRC Cluster Services
There are many practical considerations associated with running ADCIRC locally, and we are prepared to discuss all of them with you. These include:
- The advantages and disadvantages of a local physical cluster versus cloud computing, and when to use either approach
- The advantages and disadvantages of multi-core workstations or laptops compared with a dedicated cluster
- Electrical capacity, power consumption, battery backup, cooling, and additional HVAC requirements
- Shared storage, network throughput, MPI communication, and ADCIRC scaling
- Equipment cost, total cost of ownership, maintenance, and depreciation
- When to purchase used professional equipment rather than new hardware
- How to build an affordable initial cluster while retaining the ability to expand it later
We can also provide complete ADCIRC cluster services, including:
- Cluster design and detailed equipment recommendations
- Cluster hardware installation and setup
- Linux installation and system configuration
- Network and shared-storage configuration
- Slurm workload-manager installation and configuration
- Compiler, MPI, NetCDF, HDF5, and supporting-library installation
- ADCIRC and ADCIRC+SWAN installation, compilation, and testing
- ADCIRC cluster user training and IT support
- ADCIRC engineering and modeling support
- General cluster maintenance, monitoring, troubleshooting, and expansion
- Rental access to part or all of Cephas
Modern vs. 2006: What a 160-Core, ~7.0 TFLOPS Cluster Means
TL;DR: The theoretical computing capacity provided today by four professional workstations could have required several hundred dual-socket servers and a small datacenter in 2006.
| Metric |
Cephas Today |
Approximate 2006 Equivalent |
| Peak FP64 |
Approximately 7.0 TFLOPS |
Approximately 7.0 TFLOPS |
| Compute Nodes |
4 tower workstations |
Approximately 350–440 servers |
| Physical CPU Cores |
160 |
Approximately 1,400–1,760 |
| Installed Memory |
128 GB |
Approximately 1.4–3.5 TB across the equivalent servers |
| Estimated Heavy-Load Power |
Approximately 1.6–2.0 kW |
Approximately 140–176 kW |
| Physical Footprint |
4 tower workstations |
Approximately 9–11 racks at one server per rack unit, before supporting equipment |
Cephas Today
- Nodes: 4 Dell Precision T7810 workstations
- Processors: 8 × Intel Xeon E5-2698 v4
- Physical cores: 20 per processor, 40 per node, and 160 cluster-wide
- Estimated peak FP64: Approximately 0.88 TFLOPS per processor and approximately 7.0 TFLOPS total
- Installed memory: 32 GB per compute node and 128 GB total
- Estimated heavy-load power: Approximately 400–500 watts per node, or approximately 1.6–2.0 kW total
- Footprint: Four professional tower workstations
- Network: Private switched 10 GbE fabric for Intel MPI and NFS traffic
- Storage: 16 TB of NFS shared storage carried over the 10 GbE fabric
- Representative ADCIRC performance: EC2001, 254,565 vertices, 159 computational ranks plus one writer, 1,000 one-second steps in approximately 8.06 seconds
Approximate 2006 Equivalent
- Representative processors: Dual-socket, dual-core Intel Xeon “Woodcrest” or comparable AMD Opteron systems
- Estimated FP64 performance per node: Approximately 16–20 GFLOPS
- Nodes required to approach 7.0 TFLOPS: Approximately 350–440
- Physical cores: Four per node, or approximately 1,400–1,760 cores
- Typical memory: Approximately 4–8 GB per node, or approximately 1.4–3.5 TB distributed across the cluster
- Estimated server power: Approximately 400 watts per node, or approximately 140–176 kW, before cooling and other datacenter overhead
- Estimated footprint: Approximately 9–11 full racks at one rack unit per server, with additional space required for networking, storage, power distribution, and management systems
Notes:
FLOPS values are estimated theoretical peak double-precision performance figures and are intended only as rough comparisons. The 2006 estimates assume approximately 16–20 GFLOPS per dual-socket, dual-core server. Real ADCIRC performance depends on processor frequency under sustained load, memory bandwidth, network latency and throughput, compiler and MPI configuration, file-system performance, model characteristics, and application scalability. Cephas now uses a private 10 GbE fabric for MPI and NFS traffic; the current hard-disk storage subsystem may still limit some file-transfer and output workloads before the network reaches its maximum throughput.
|
|