About the job
Job Description
The incumbent will be responsible for the architecture, design, implementation and day-to-day operations of the SIA Group's core Data Centre network, keeping mission-critical services highly available, performant, and secure. He/She requires an analytical mindset, good awareness and appreciation of operational practices, and strong expertise in automation and AI-enabled operations in order to reduce toil, improve incident response, and strengthen resilience at scale.
Key Key Responsibilities
- Service Operations & Reliability
*Provide operational support for core DC networking (routing/switching, segmentation, connectivity) *Drive Incident, Problem, Change, and Configuration MANAGEMENT to meet service targets and standards *Serve as technical lead during major incidents, performing deep-dive troubleshooting, RCA, and driving corrective/preventive actions *Plan and execute complex changes with risk assessment, readiness checks, back-out planning, and post-change validation *Maintain high-quality configuration and operational documentation; continuously improve runbooks/SOPs
- Operational Automation & AI-enabled Ops
*Design and build automation to improve BAU consistency and efficiency (e.g., config deployment/validation, compliance, health checks, reporting) *Use Ansible, Python, and APIs to standardise operational workflows.
*Apply AIOps/AI-assisted operations (event correlation, anomaly detection, noise reduction, and predictive alerting) to reduce MTTR *Scale automation patterns
- Security Operations & Zero Trust
*Operate and improve the security posture of DC network infrastructure *Support micro-segmentation and Zero Trust-aligned controls within the data centre network *Proactively identify resiliency and security gaps and drive remediation through engineering changes and automation
- Vendor & Service Delivery Management
*Collaborate with vendors and stakeholders to ensure service availability and timely issue resolution *Track incidents/problems/tickets to ensure SLA adherence, quality updates, and timely closure *Participate in service delivery reviews and drive actions that improve outcomes
- Continuous Improvement & Technology Refresh (Ops-led)
*Identify operational pain points and implement improvements across process, tooling, monitoring, and automation *Participate in POCs/lab validations to improve operability, reliability, and automation
Requirements
*Bachelor's degree in Computer Sciences or related discipline *6+ years of hands-on network operations/engineering experience in a multi-vendor environment.
*Strong DC networking fundamentals (routing/switching, troubleshooting, complex change execution).
*Experience with operational practices: incident/problem/change management, RCA, and service reliability.
*Experience implementing process re-engineering and automation to improve processes.
*Strong problem-solving skills with the ability to isolate and resolve issues in complex environments.
*Excellent team player; ability to manage conflicts to achieve common goals.
*Strong communication skills; comfortable engaging technical and non-technical stakeholders.
*Experience in two or more of: core DC networking, security, network automation.
*Working knowledge of Ansible/Python/APIs (or motivation to build this capability quickly).
*Familiarity with segmentation/micro-segmentation and operational security hygiene.
*Certifications such as Cisco/F5/Palo Alto/Check Point (or equivalent).
*Awareness of AIOps/AI trends and ability to apply them pragmatically.
We thank all candidates for your interest in Singapore Airlines, and regret that only shortlisted candidates will be notified.
