See all the jobs at Xideral here:
, , | Software Development | Full-time | Fully remote
Seeking Experienced Senior AWS Site Reliability Engineers for Exciting Projects – Remote in Mexico
We are looking for skilled Senior Site Reliability Engineers with a minimum of 5 years of experience to join a dynamic team within a leading organization. This role involves supporting and improving cloud operations for microservice-based platforms, with a focus on production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.
Key Responsibilities:
- Own and improve the reliability of cloud-based services and supporting infrastructure.
- Participate in on-call rotations and support production systems outside normal business hours.
- Lead incident response activities, including triage, escalation, mitigation, and service restoration.
- Drive blameless postmortems and ensure corrective actions are tracked to closure.
- Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis.
- Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
- Support and improve cloud and container platforms across AWS and Azure.
- Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
- Build automation to reduce manual effort and improve operational efficiency.
- Configure and improve monitoring, alerting, logging, diagnostics, and observability.
Technical Skills Required:
With over 5 years of experience as a Senior Site Reliability Engineer, you must be proficient in the following technical skills:
- Strong hands-on experience with AWS and Azure cloud platforms.
- Strong experience with Terraform for Infrastructure as Code (IaC).
- Experience with Atlantis, ArgoCD, or similar infrastructure and deployment automation tools.
- Strong hands-on experience with Docker and Kubernetes.
- Experience designing, maintaining, and troubleshooting complex CI/CD pipelines.
- Strong production support experience, including incident management, Root Cause Analysis (RCA), postmortems, and runbook creation.
- Strong observability experience, including monitoring, alerting, logging, diagnostics, and performance analysis.
- Good understanding of cloud networking, security, access controls, and InfoSec practices.
- Experience with version control, branching, merging, pull requests, and conflict resolution.
- Understanding of cloud cost optimization and resource utilization.
Good-to-Have Skills:
- Experience with microservice-based platforms.
- Experience with Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics, or similar tools.
- Scripting or programming experience using Python, Bash, Go, or Java.
- Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
- Experience with disaster recovery testing and production readiness reviews.
- Prior experience mentoring junior engineers or leading technical troubleshooting.
Qualifications:
- Bachelor’s degree or higher.
- Fluent in English (Advanced).
- Excellent communication, empathy, commitment, leadership, teamwork, and a proactive attitude.
Location & Schedule:
- Remote work from Mexico.
- Preferred hybrid model in Guadalajara, Jalisco, with expected onsite attendance 2 days per week.
- Work hours Monday to Friday, 09:00 – 18:00.
- Advanced English skills are mandatory, and only residents of Mexico.
Benefits:
- Attractive Salary + Premium Benefits
- Performance bonuses, grocery coupons, and savings are found.
- Aguinaldo, premium vacations, and vacations paid
- SGMM Medical insurance, family, and Life insurance.
Candidates must include their compensation expectations in their applications and resumes in English.
Interested? Apply now through this link:
Fetching your Linkedin profile ...