Site Reliability Engineer
- Xiaomi
- Singapore, Singapore
- SGD 90,000 – SGD 140,000
新加坡社招全职职位 ID:A06044
职位描述
- Job Responsibilities
- Ensure the stability, reliability, and efficient operation of the Xiaomi's global business, maintaining high availability of services at all times.
- Responsible for core operational tasks such as resource provisioning and management, incident response, capacity management, monitoring, and reliability improvements.
- Review technical architecture design, assess soundness of the design, and proactively identify and resolve reliability risks.
- Conduct in-depth analysis of systemic deficiencies, identify bottlenecks and develop optimization strategies; plan and execute projects to improve system reliability and ensure cost-effectiveness and highly availability of the systems.
- Participate in 24/7 on-call rotation, promptly respond to and resolve production incidents to ensure service availability.
- Analyze and improve processes to build stable, highly available systems; drive continuous automation improvements, and minimize manual intervention.
职位要求
- Job Requirements
- Bachelor’s degree in Computer Science or a related field.
- Proficiency in one of the following programming languages: Python, Go, or shell scripting, with demonstrated ability to independently develop modules or platforms.
- Familiar with cloud computing; experience in managing multi-cloud or hybrid cloud platforms (e.g., Alibaba Cloud, Azure, AWS) is preferred.
- Strong foundation in computer science, with hands-on experience in Linux, networking, load balancing, and designing high-availability and disaster recovery architectures.
- A good team player with a strong sense of responsibility, self-driven and highly motivated.
- Minimum 3 years of working experience in operations and maintenance of large-scale web services is preferred; hands-on experience in managing or operating large-scale web services or projects is a plus.
- Fluent in Mandarin (spoken) is a plus.
投递
Skills
- Python
- Go
- Shell Scripting
- Cloud Computing
- Multi-cloud Management
- Incident Response
- Capacity Management







