100% remote work
Job Description
This position is needed to evolve and maintain fundamental Compute infrastructure, collaborating with a passionate team to enhance the system capabilities. Key responsibilities include develop scalable cloud-native environments, VM orchestration, AWS-ASG Auto Scaling group of EC2 instances, hardened base AMIs, and secure container images while maintaining critical OS libraries. You'll implement solutions for providing robust legacy platform support and cutting-edge cloud-native technologies, all driven by automation and best practices.
Join us in shaping the future of our Compute infrastructure
- Collaborate with Tech Leaders, Architects and other Engineers to develop solutions for complex problems in distributed computing and infrastructure management.
- Automate solutions for operational issues, such as monitoring, performance, planning, and disaster response.
- Participate in an on-call rotation to support our business-critical infrastructure.
- Ensure a high quality implementation by applying Infrastructure as Code industry standards.
- Demonstrate effective communication by authoring and reviewing design documents, runbooks, and other service documentation, and keeping a good record of changes in the systems.
- Apply Agile methodologies to continuously deliver value to the customers.
- Act as point of contact for legacy/new Compute system components.
Requirements
- 2+ years of experience in AWS Cloud infrastructure management. (preferably backend/infrastructure-focused like AMI, EC2, IAM policies/roles, etc.).
- Strong ASG (Auto Scaling Groups) knowledge, to design, implement, and support scalable cloud-native environments.
- Experience with hardened base AMIs and AL23.
- Proficiency with one or more programming languages: like Java or Python (includes SW Arch patterns, clean code, debugging, etc)
- Proficient in shell scripting to streamline repetitive tasks and enhance efficiency in operations.
- Skills to work independently with multiple global teams, developing, configuring, deploying, and operating the global Twilio Infrastructure Platform, blending operational excellence with development best practices.
- Knowledge of container-based application/services.
Nice to have
- Knowledge on deployment tools and frameworks like infrastructure as a code and continuous deployment processes (ex: Github, Buildkite, Terraform-TFC, ArgoCD, Harness, Cloud network).
- Operational experience in complex distributed systems, including experience with SLO/SLAs towards high availability and reliability goals, including tools like DataDog or Phometheus.
- Exposure to File Integrity Monitoring (FIM) tools, specifically Falco, and awareness of compliance frameworks (PCI, SOX).
- Knowledge in Kubernetes
- Experience with Claude AI or similar.
Company offers
- Working at Twilio offers many benefits, including competitive pay, generous time off, ample parental and wellness leave, healthcare, a retirement savings program, and much more. Offerings vary by location.