Site Reliability Engineer
24 Jul 2026Job Description:
We are looking for a Site Reliability Engineer to work on a enterprise project.
This role requires someone who has strong experience in Kubernetes platform.
Job Responsibilities
- Monitor, analyze, and optimize system, application, and infrastructure performance using observability and monitoring tools to ensure platform reliability and availability.
- Design, implement, and maintain reliable, scalable, and secure infrastructure platforms supporting enterprise applications and services.
- Administer and support Linux, Windows, virtualization, and Kubernetes environments to ensure high availability and operational excellence.
- Collaborate with development, infrastructure, and operations teams to improve platform reliability through automation, standardization, and best practices.
- Participate in platform architecture reviews, system design, infrastructure planning, and capacity management to support business growth.
- Develop and maintain Infrastructure as Code (IaC), automation scripts, and operational workflows to improve efficiency and reduce manual effort.
- Build and enhance CI/CD pipelines to support reliable and efficient application deployment and infrastructure provisioning.
- Proactively identify system issues, performance bottlenecks, and operational risks, implementing preventive measures and automation to improve service resilience.
- Define, measure, and maintain Service Level Objectives (SLOs), Service Level Indicators (SLIs), and operational metrics to ensure service reliability.
- Perform incident response, root cause analysis, and post-incident reviews to continuously improve platform stability and operational processes.
- Work closely with project and delivery teams to ensure infrastructure services are delivered on schedule while maintaining operational stability and service quality.
- Develop and maintain technical documentation, operational procedures, and knowledge articles to support ongoing platform operations.
Job Requirements
- Bachelor's Degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Minimum 5 years of experience in IT infrastructure operations, cloud platforms, Site Reliability Engineering (SRE), DevOps, or related fields.
- Strong understanding of Linux and Windows server administration in enterprise environments.
- Experience with container technologies and orchestration platforms, including Docker and Kubernetes.
- Hands-on experience with infrastructure automation, scripting, or programming using one or more high-level programming languages (e.g., Python, Go, Java, PowerShell, or Bash).
- Good understanding of CI/CD pipelines, automation tools, and modern software delivery practices.
- Experience with infrastructure management, system monitoring, and operational support.
- Knowledge of VMware vSphere/VMware Cloud Foundation (VCF) or Broadcom VMware technologies is an advantage.
- Excellent written and verbal communication skills with the ability to collaborate effectively across cross-functional teams.
- Strong analytical, troubleshooting, and problem-solving skills with attention to detail.
- Self-motivated, customer-focused, and able to work independently while managing multiple priorities in a fast-paced environment.
Working location at the East region of Singapore.
This role requires individual to clear CAT 1 due to the nature of the project.
We regret that only shortlisted candidates will be notified.
Interested applicants kindly click on apply now or send your updated resume to ••••@peopleprofilers.com
Contact No. : +65 •••• •731
Berlyn Lum Miao Yu
Registration Number: R1766577
EA License Number: 02C4944
People Profilers Pte Ltd, 20 Cecil St, #08-09, PLUS Building, Singapore 049705
http://www.peopleprofilers.com
Salary:
S$ 6,000.00
-
S$ 7,800.00
/