Network Engineer/Architect – NVIDIA AI Cloud | ID:335
- Full Time
- €4000 - €5500 / Month
Job Description
The platform connects thousands of NVIDIA Blackwell GPUs (DGX B200 and RTX Pro Servers) into a single low-latency, high-throughput network for European industry, research, and the public sector. You would join a stable, large technology organization with deep infrastructure and operations experience, on a sovereign cloud held to strict security and availability standards. For a network engineer who wants to work on AI-scale fabric instead of standard enterprise networking.
Role overview:
You design, build, automate, and operate the network platform behind the AI cloud: the InfiniBand and RoCE fabric, switches, firewalls, routers, and border gateways that form the core interconnect. You bring the network up through provisioning, keep it running across its lifecycle, and tune it for the throughput that GPU workloads demand. You come in as a senior network engineer, hands-on with operations and automation, and you grow from there into network architecture: owning the design of the fabric and the standards the rest of the team builds on. How fast that happens is up to you. You work across infrastructure, platform, and AI teams under ITIL processes, with 24/7 operations and an on-call rotation, and you are a technical point of contact for customers running on the platform.
Key Responsibilities:
- Design and build the network platform for the AI cloud: InfiniBand and RoCE fabric, switches, firewalls, routers, and border gateways.
- Provision and maintain InfiniBand switches and the unified fabric (InfiniBand plus Ethernet / RoCE) at scale.
- Automate configuration, deployment, and lifecycle management with scripting and Infrastructure as Code (Ansible, Terraform, Helm, SaltStack).
- Manage firewalls and network security: FortiGate policies, NAT, VPNs, IDS/IPS, HA, segmentation, DDoS mitigation, and zero-trust networking.
- Run data center routing and border connectivity: BGP/OSPF, ASNs, IP transit, peering, and failover.
- Handle OS and firmware across the network estate: patches, upgrades, configuration backups, and critical patch management.
- Set up monitoring, observability, and root-cause analysis for a large-scale data center network (Prometheus, Grafana, UFM, NVIDIA/Mellanox diagnostics).
- Follow and improve ITIL incident, problem, and change workflows, document runbooks, and hold to zero-outage standards.
- Take part in 24/7 operations and on-call rotation, with independent troubleshooting of incidents and hardware cases.
- Work with platform and AI teams on deployments, and advise customers on the network side of their workloads.
- Grow into owning the network design, the standards, and the architecture the team builds and operates against.
Requirements
- 5+ years in network engineering for large-scale or data center networks, with the drive to grow into network architecture.
- Deep understanding of InfiniBand architecture, RoCE, and low-latency, high-throughput networking for AI or HPC workloads.
- Hands-on NVIDIA/Mellanox switch configuration and UFM (Unified Fabric Manager).
- Data center routing: BGP/OSPF, ASNs, IP transit, peering, failover, on Cisco or Juniper routers.
- Strong Linux networking (Cumulus OS, Ubuntu, Debian): bridges, bonds, VLANs, routing tables.
- FortiGate firewall administration and network security: segmentation, DDoS mitigation, zero-trust.
- Configuration and lifecycle management: provisioning, firmware and OS upgrades, patching, backups.
- Scripting (Go, Python, or Bash) and automation (Ansible, Terraform, Helm, SaltStack).
- ITIL processes (incident, problem, change), plus familiarity with NOC/SOC and on-call models.
- English at C1, used actively with customers and across teams.
A strong plus:
- German (active), especially for work with German-speaking customers.
- Diagnostic tooling: iperf, ethtool, nvidia-smi, perfquery.
- Kubernetes, VMware Tanzu Kubernetes, CI/CD in Kubernetes, Git-based automation.
- Software-Defined Networking (SDN) and NVIDIA GPU-accelerated server platforms.
Required Education
- University degree in IT, computer science, or a related field. Strong hands-on experience counts more than a specific diploma, so we look at what you have built, not only at your formal degree.
Required Language
- English - C1
- German - advantage
Suitable For Graduates
No
Skills
Employee Benefits
- Work on the network fabric of one of Europe’s largest AI factories, on InfiniBand and RoCE at production scale.
- Direct work with around 10,000 NVIDIA GPUs and the high-speed fabric that connects them, hardware most network engineers never get near.
- A cloud platform developed in-house, without vendor lock-in. You work on real network engineering, not on managed-service configuration screens.
- Strong focus on security, compliance, and sensitive data, with customers in banking, insurance, and defense.
- A clear path from senior network engineer to network architect, on a live AI-scale fabric, with mentoring and real ownership as you grow.
- Work alongside top infrastructure engineers who design and run this fabric every day.
- An internal reskilling program, training, and a learning budget, including InfiniBand and NVIDIA fabric know-how that is hard to get anywhere else.
- Hybrid or full remote within Slovakia, with occasional travel to the data center or customers.
- A stable employer with a large engineering base, so the project has real backing and the role has runway.
- Over 25 company benefits across finance, health and sport, learning and development, and family and work-life balance.