NVIDIA AI Cluster Delivery Engineer
Turning GPU servers, cluster networks, and system software into acceptance-ready AI infrastructure.
Focused on GPU cluster deployment, network integration, system delivery, acceptance testing, and operations handover.
Home views: 44
Selected Delivery Work
Anonymized Project Cases
Proof through delivery scenarios: scale, responsibility, difficulty, and result.
AI Datacenter GPU Cluster Delivery
- Scale
- Dozens of GPU servers with IB / RoCE high-speed network
- Role
- Handled OS deployment, driver installation, network integration, NCCL validation, and acceptance documents.
- Result
- Completed acceptance and delivered deployment records, issue lists, and handover materials.
AI Training Cluster Performance Acceptance
- Scale
- Multi-node multi-GPU training environment
- Role
- Checked consistency across GPU, network, storage, OS, and software versions.
- Result
- Identified performance fluctuation causes and formed a reusable acceptance workflow.
GPU Node Troubleshooting And Recovery
- Scale
- Pre-production delivery site
- Role
- Resolved GPU disappearance, driver errors, unstable links, and compatibility issues.
- Result
- Shortened issue closure time and protected the acceptance schedule.
Capability Matrix
AI Cluster Delivery Capability Matrix
Covering hardware, systems, networking, drivers, schedulers, validation, and handover.
GPU Servers And Systems
Driver And Compute Stack
Cluster Network
Scheduling And Platform
Delivery Playbook
From Site Arrival To Acceptance Handover
Recruiters need to see not only tools, but the ability to drive real delivery.
- 01 Requirements
- 02 Rack
- 03 Cable
- 04 Install OS
- 05 Deploy Drivers
- 06 Integrate
- 07 Validate
- 08 Handover
Technical Knowledge note
Cluster Delivery Technical Notes
Structured SOP manuals, troubleshooting playbooks, and hardware test.py records.
Contact
Focus on NVIDIA AI cluster delivery, AI datacenter, and GPU infrastructure.
China
JnZhang0810@163.com