FULL_TIME
Asset & Spare Parts Manager (Data Centre Management)
The Asset & Spare Parts Manager oversees the lifecycle, inventory, availability, and documentation of critical data centre assets and spare parts. This role ensures hardware assets are accurately recorded, critical spare parts stay available to meet operational and service-level requirements, and asset, warranty, logistics, and lifecycle activities are properly controlled and documented.
Key Responsibilities:
- Operate and support high-performance network fabrics interconnecting GPU compute infrastructure, including InfiniBand and/or RoCE-based environments.
- Oversee network fabric topology, fabric management functions, congestion control, link health, and network performance.
- Track network health indicators and link error conditions, detect potential issues, and drive proactive remediation as needed.
- Support data centre network architectures encompassing spine-leaf fabrics, overlay networks, routing, and management network segmentation.
- Work with technologies and environments including EVPN/VXLAN, BGP, OSPF, and related data centre networking technologies.
- Handle network incidents, changes, and problem resolution across the GPU platform, including maintenance windows and change governance processes.
- Carry out network equipment firmware upgrades, configuration tasks, and hardware or connectivity validation.
- Verify network cabling, optics, DAC/AOC connections, and other physical network components to maintain reliable connectivity.
- Coordinate with equipment vendors, OEMs, and technical support teams regarding hardware failures, complex incidents, and infrastructure maintenance.
- Support implementation and review of network security policies, segmentation, and access controls in line with applicable standards.
- Monitor and report network fabric health, utilisation, performance, and capacity.
- Contribute to capacity planning and provide technical recommendations for future GPU cluster expansion and network scale-out.
- Maintain operational documentation, network records, incident/change records, and compliance evidence.
- Identify opportunities for network optimisation, automation, reliability improvements, and operational efficiency.
Must-Have Requirements:
- 7–11 years of experience in data centre networking, network engineering, network operations, or related infrastructure environments.
- Strong hands-on experience with high-performance data centre network fabrics supporting GPU, HPC, or similarly demanding compute environments.
- Experience with InfiniBand and/or RoCE-based networking, including fabric operation, monitoring, performance management, and troubleshooting.
- Experience with fabric management, topology management, congestion control, and network performance optimisation.
- Strong understanding of data centre switching and network fabric architecture, including spine-leaf designs.
- Experience with network overlays and data centre technologies such as EVPN/VXLAN.
- Strong knowledge of routing protocols including BGP and OSPF, with exposure to WAN/network connectivity technologies such as MPLS.
- Experience with network segmentation, management networks, and access control implementation.
- Experience with network security and firewall policy management, including enterprise firewall environments.
- Strong packet-level troubleshooting capabilities, including analysis of connectivity, performance, and network errors.
- Hands-on understanding of data centre physical networking, including structured cabling, optics, DAC/AOC connections, and related connectivity standards.
- Experience performing network equipment configuration, firmware upgrades, validation, and troubleshooting.
- Exposure to network automation and configuration management using tools or scripting such as Python and Ansible.
- Experience handling network incidents, changes, problems, maintenance activities, and technical escalations.
- Ability to work within an 8x5 operational environment with on-call support as required.
- Good communication skills in Bahasa Indonesia and working English.
Preferred Qualifications:
- Experience supporting GPU, AI, machine learning, HPC, or high-density data centre environments.
- Experience with NVIDIA-based GPU networking and high-performance interconnect technologies.
- Experience with network fabric management and monitoring platforms such as NVIDIA UFM or equivalent.
- Exposure to NVIDIA Quantum InfiniBand platforms, including NDR/HDR environments.
- Experience with RoCE v2 and performance or congestion optimisation for high-performance workloads.
- Experience with data centre networking platforms such as Cisco ACI, Juniper QFX/EX, NVIDIA networking platforms, or equivalent technologies.
- Experience with network management and telemetry platforms such as NVIDIA Cumulus/NetQ or equivalent.
- Experience with firewall platforms such as Palo Alto or Fortinet.
- Experience coordinating with OEM technical support/TAC teams for network hardware incidents.
- Experience with network capacity planning and infrastructure design for large-scale cluster expansion.
- Familiarity with IT service management, change governance, and maintenance processes.
- Relevant professional certifications such as NVIDIA NCP-AII / NCA-AIIO, CCIE Enterprise Infrastructure, CCNP Data Center, Juniper JNCIP-DC, Palo Alto PCNSE.
Key Attributes:
- Strong network troubleshooting and analytical skills, especially at packet and infrastructure level.
- Good understanding of high-performance and data centre networking environments.
- Strong operational awareness with the ability to proactively identify network risks and potential service-impacting conditions.
- Detail-oriented in network configuration, cabling, documentation, monitoring, and change activities.
- Able to manage network incidents and technical issues through structured troubleshooting and escalation processes.
- Strong understanding of network security, segmentation, and access control principles.
- Comfortable working with cross-functional teams, infrastructure teams, vendors, OEMs, and technical stakeholders.
- Able to balance operational stability with network optimisation, automation, and infrastructure scalability.
- Strong documentation and reporting discipline.
- Comfortable supporting maintenance activities and on-call operational requirements.
Asset & Spare Parts Manager (Data Centre Management)
PT Leap Digital Indonesia • Jakarta • Gaji tidak dicantumkan
Deskripsi Pekerjaan
Ringkasan
- Perusahaan
- PT Leap Digital Indonesia
- Lokasi
- Jakarta
- Tipe kerja
- FULL_TIME
- Gaji
- Gaji tidak dicantumkan
- Tanggal tayang
- 29 Sep 2026
- Terlihat sejak
- 2 Okt 2026
Sumber & keterlacakan data
JobStreet / Jobsdb (SEEK) — https://id.jobstreet.com/job/94948405
Siap melamar? Langsung di sumber asli.Buka Halaman Lamaran
Lowongan Terkait Lainnya
FULL_TIME1mg
L1 Data Centre Operator
PT Leap Digital Indonesia • Jakarta
Gaji tidak dicantumkanLihat Detail →
FULL_TIME1mgManageEngine Integration Developer
PT Leap Digital Indonesia • Jakarta
Gaji tidak dicantumkanLihat Detail →
FULL_TIMEHari IniEstimator
PT Takenaka Indonesia • Jakarta
Gaji tidak dicantumkanLihat Detail →
FULL_TIMEHari IniArea Sales Representatives (Jakarta)
PT. Indomak Kitacipta Karya • Jakarta
Rp 5.700.000 – Rp 8.000.000 per monthLihat Detail →