Job Requirements
Remote Arlington, VA
Secret Polygraph not specified
Mid Level Career (5+ yrs experience)
$180,000 - $200,000
Job Description
Site Reliability Engineer:
Our client is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is based in Arlington, VA, as a hybrid/remote position.
Responsibilities:
-Design and maintain highly available production systems.
-Define and manage SLIs, SLOs, and error budgets.
-Automate operational tasks and eliminate manual processes.
-Develop monitoring, alerting, and observability solutions.
-Improve system performance, capacity, and resilience.
-Lead incident response and root cause analysis.
-Implement disaster recovery and continuity strategies.
-Partner with development teams to improve application reliability.
Required Skills and Experience:
-Bachelor's with 12+ years of infrastructure/cloud engineering experience (or commensurate experience)
-5–10+ years of engineering experience, with a strong background in Linux and Windows systems
-Expertise in Kubernetes and container platforms
-Experience working with cloud infrastructure environments
-Proficiency in scripting languages such as Python and Go
-Hands-on knowledge of Terraform and automation tools
-Familiarity with monitoring platforms and incident management practices
-Experience designing and managing CI/CD pipelines
Preferred Skills and Experience:
-Kubernetes certifications
-AWS/Azure certifications
-DevOps certifications
-ITIL preferred
Our client is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is based in Arlington, VA, as a hybrid/remote position.
Responsibilities:
-Design and maintain highly available production systems.
-Define and manage SLIs, SLOs, and error budgets.
-Automate operational tasks and eliminate manual processes.
-Develop monitoring, alerting, and observability solutions.
-Improve system performance, capacity, and resilience.
-Lead incident response and root cause analysis.
-Implement disaster recovery and continuity strategies.
-Partner with development teams to improve application reliability.
Required Skills and Experience:
-Bachelor's with 12+ years of infrastructure/cloud engineering experience (or commensurate experience)
-5–10+ years of engineering experience, with a strong background in Linux and Windows systems
-Expertise in Kubernetes and container platforms
-Experience working with cloud infrastructure environments
-Proficiency in scripting languages such as Python and Go
-Hands-on knowledge of Terraform and automation tools
-Familiarity with monitoring platforms and incident management practices
-Experience designing and managing CI/CD pipelines
Preferred Skills and Experience:
-Kubernetes certifications
-AWS/Azure certifications
-DevOps certifications
-ITIL preferred
group id: 10112344
Defining Company Culture