Senior Site Reliability Engineer
andglobal · Ulaanbaatar
Job description
About the role
As a Senior Site Reliability Engineer you will design, build and maintain highly available, secure and scalable infrastructure for AND Global and its subsidiaries. You will work closely with engineering teams to automate operations, ensure system stability and support new business initiatives.
Key responsibilities
- Monitor production systems, application availability, performance and overall health.
- Build and maintain reliable cloud and Kubernetes infrastructure.
- Automate provisioning, deployment, monitoring and recovery processes.
- Investigate production incidents, identify root causes and implement permanent fixes.
- Lead incident response and prepare post‑incident reports.
- Collaborate with development teams to improve application reliability and release processes.
- Define and maintain service‑level indicators, objectives and availability targets.
- Improve CI/CD pipelines and deployment procedures.
- Perform capacity planning, performance tuning and cost optimization.
- Build dashboards, alerts, monitoring standards and operational runbooks.
- Maintain backup, disaster recovery and high‑availability procedures.
- Identify technical risks, bottlenecks and areas for continuous improvement.
- Mentor junior engineers and promote engineering best practices.
Required profile
- Bachelor’s degree in computer science, IT or equivalent practical experience.
- At least 5 years of experience in SRE, DevOps, Platform or Cloud engineering.
- Strong Linux administration and troubleshooting skills.
- Production experience with AWS, Microsoft Azure or Google Cloud.
- Extensive experience with Kubernetes and Docker.
Required skills
- Infrastructure‑as‑code tools such as Terraform or OpenTofu.
- CI/CD tools like GitLab CI, GitHub Actions, Jenkins or Argo CD.
- Monitoring and observability tools such as Prometheus, Grafana, Loki, ELK, OpenSearch, Datadog or OpenTelemetry.
- Scripting or programming in Python, Go or Bash.
- Networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancers, VPNs, firewalls and routing.
- Support for distributed applications, databases, storage and messaging systems.
- Incident management, root‑cause analysis and post‑mortem reporting.
Questions fréquentes
Why are you reporting this job?
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 3 цагийн өмнө
Expires одоогоос 1 сарын дараа
3 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
andglobal
Ulaanbaatar