Thermal-paste and memory-pad replacement, alongside improved airflow, lowered memory temperatures by approximately 10–20 °C on the RTX 3080 and 3090 cards. Sustained testing exposed unstable settings before unattended operation. Each GPU had workload-specific profiles, with higher core clocks for gaming and rendering and higher memory clocks for memory-intensive workloads. Underclocking and undervolting reduced unnecessary core power consumption. In one recorded tuning session, throughput increased by 20.4% while GPU power rose by 5.2%, improving throughput per watt by 14.4%.
The setup expanded incrementally as GPUs were added. Open-frame construction provided access for maintenance and upgrades: one rig could be shut down to replace or add a card while the remaining machines continued operating. Initially operated nearby for supervision and troubleshooting, the rigs later moved into a separate network-connected building. Headless access, remote monitoring and automatic recovery from network interruptions, power loss and updates supported continuous operation, including extended periods overseas without physical access.


