This role has closed just now. It is no longer on NVIDIA's board, so there is nothing left to apply to. The posting is kept here because you saved it or opened it; it is a record, not an offer.

Verified by our engine · Posted 4w ago

NVIDIA

Senior Software Engineer, Network System Validation

NVIDIA · Engineer, Sys SW

Switzerland

Senior

SoftwareAnsibleBash / ShellC++Distributed systemsKubernetesLLM / GenAIPythonRust

Last seen 4h ago

Posted
4w ago

Posted on 9 September 2026

Workplace
Remote

Work model: Remote

Salary
Not disclosed

Salary range not shared by the company

Visa sponsorship
Not specified

Visa sponsorship details unknown

This role has closed. It's kept as a record — see NVIDIA's open roles or the similar live roles below.

See NVIDIA's open roles

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Within NVIDIA, the Networking Business Unit (NBU) builds the high-speed interconnect — Ethernet, InfiniBand, NVLink, and BlueField DPUs — that switches thousands of GPUs into a single AI supercomputer, moving data at the scale and speed the most demanding workloads require.

We are looking for a Technical Lead to join our Network System Validation group. You will work on validating advanced networking solutions across NVIDIA complex AI cluster environments. The group is a high-performance engineering force that treats validation as a first-class software problem. We build systems, frameworks, and benchmarks that prove our network's correctness and performance at scale. In this role you will lead the validation direction and engineering excellence of one of our technology validation teams. This is a deeply hands-on role for a technology leader who can own the technical roadmap, technology mentoring a team of high-performance engineers, and push NVIDIA's network to its speed-of-light limits. The role combines the development of validation methodologies and automation tools with hands-on debugging, performance analysis, and investigation of cutting-edge AI networking technologies at scale.

What you’ll be doing:

Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions

Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.

Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution

Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities

Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection

Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations

Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes

What we need to see:

B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience

8+ years of experience in networking, system validation, or related domains

Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause

Ability to read, debug, and reason about C/C++ code (Rust or Go a plus)

Strong scripting and automation experience using Python, Bash, and/or Ansible

Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress

Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed

Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines

Ways to stand out from the crowd:

Experience with large-scale clusters or distributed systems

Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField)

Background in performance analysis, Kubernetes, or cloud environments

Background in chaos testing, fault injection, or simulation systems

We have some of the most forward-thinking and hardworking people working for us. If you're creative and autonomous, we want to hear from you!

About NVIDIA

NVIDIA's European sites — Munich, Zurich, Kyiv, Amsterdam — cover deep learning, driver and systems software rather than sales engineering. One of the few places in Europe doing serious GPU and inference infrastructure work.

Similar roles

  1. PRO members only Pro Posted today
  2. Solution Manager - Architect (with Avaloq Experience) New Hybrid Avaloq Sales & Delivery Zurich Software High-growth Not disclosed Visa: Not specified yesterday
  3. Software Engineer, Robot Autonomy, Intern New On-site Laelaps AI Engineering Zurich SoftwareDockerGitLinux +1 High-growth Not disclosed Visa: Not specified yesterday
  4. Software Engineer III, Merchant Shopping, Intelligence and Agents New Google Google Zurich SoftwareC++JavaKotlin +2 Top-tier Not disclosed Visa: Not specified yesterday
  5. Durable Workflows Engineer New Swisscom Data Administration & Analysis Zurich SoftwareAWSCI/CDDistributed systems +3 High-growth Not disclosed Visa: Not specified yesterday
  6. PRO members only Pro Posted today
3,766 more roles like this. Open the board →