careertakescareertakes

Software Engineer - SRE, Retail & Pharmacy

Confidential ClientHealthcare
Entry LevelFull-timeOn-SiteN/A
Woonsocket, Rhode Island
3 Days Ago

The Software Engineer - SRE at Confidential Client is responsible for building and maintaining the reliability, availability, and performance of distributed retail and pharmacy systems across thousands of locations. This role involves implementing observability, automating operational tasks, responding to incidents, and contributing to continuous reliability improvements in a dynamic, large-scale edge computing environment. The engineer will collaborate closely with development teams to embed reliability into services and drive operational excellence through automation and monitoring.

Boost your chances. Upload your resume to see your match score.

Unlock Match Rate
Description

About Careertakes

👉 Important disclosure: Careertakes is a third-party recruiting platform supporting this hiring process. If selected, you will be employed directly by our client, Hospitals and Health Care.

Applicants for this role may also receive access to additional matched opportunities through the Careertakes platform.


About the Team

Join an SRE practice that supports distributed retail and pharmacy systems at fleet scale. The team focuses on reliability across store Point-of-Sale, pharmacy dispensing systems, edge compute nodes in store locations, and hybrid cloud services. The operating philosophy centers on Detection, Prevention, Recovery, Learning Loops, and Developer Experience (DevX). You will help transfer reliability capability to engineering teams so they can operate independently and reduce operational toil.


What You’ll Do

Detection & Observability

  • Implement alerting and dashboards using Prometheus, Grafana, Loki, Jaeger, and OpenTelemetry to provide visibility into distributed retail and pharmacy systems.
  • Help define SLIs and SLOs for assigned services; learn how error budgets map to customer and patient impact.
  • Build and maintain dashboards for availability, latency, error rate, and saturation tied to Critical User Journeys.
  • Participate in alert tuning; document false-positive patterns and sources of noise for remediation.
  • Learn and apply anomaly-based detection and time-series baseline techniques.

Prevention & Reliability Engineering

  • Participate in Production Readiness Reviews and execute checklist items.
  • Write unit/integration tests for SRE tooling and automation.
  • Improve and maintain runbooks: identify gaps, propose fixes, and implement changes.
  • Support performance testing and reliability audits under guidance.
  • Apply reliability patterns in code and configuration (circuit breakers, retries, timeouts, bulkheads).
  • Develop a fleet-operations mindset for cohort rollouts and managing configuration drift for unattended device fleets.

Incident Response & Recovery

  • Participate in on-call rotations and escalate promptly following escalation paths.
  • Execute validated runbooks during incidents and document actions and timelines in incident-management systems.
  • Triage alerts, provide context and hypotheses, and route to domain owners.
  • Drive post-incident action items to closure.
  • Build operational familiarity with Kubernetes clusters, store servers, networking, and connectivity failure modes for edge fleets.

Learning Loops & Continuous Improvement

  • Contribute to postmortems and own follow-up monitoring, alerting, or runbook work to reduce recurrence.
  • Maintain operational documentation and review after incidents.
  • Follow a structured learning path: SLO fundamentals, incident command, chaos engineering concepts, and domain knowledge for retail/pharmacy systems.
  • Track personal MTTD and MTTR trends and use data to set improvement goals.

Developer Experience & Automation (DevX)

  • Eliminate toil by writing automation in Python, Bash, or Go; measure and document toil reduction.
  • Improve CI/CD pipelines using GitHub Actions, ArgoCD, Helm and apply GitOps practices.
  • Collaborate with development teams to instrument features for observability and reliability by design.
  • Apply containerization and cloud-native patterns with Kubernetes and infrastructure-as-code tools (Terraform/Ansible).

Required Qualifications

  • 2+ years experience in SRE, DevOps, platform engineering, or related roles with production responsibility.
  • 2+ years delivering software in large-scale distributed environments with applied reliability concepts.
  • 1+ year production-quality experience in at least one language: Python, Go, Bash, or Java.
  • 1+ year hands-on cloud experience (AWS, Azure, or GCP).
  • Practical experience with observability/monitoring tools such as Prometheus, Grafana, ELK, Splunk, Datadog, or Dynatrace.
  • Foundational understanding of containerization and orchestration with Kubernetes and Docker.
  • Strong written and verbal communication; ability to engage technical and non-technical stakeholders.
  • Comfortable contributing to and shaping processes in an environment undergoing active transformation.

Preferred Qualifications

  • Experience supporting retail, pharmacy, healthcare, or other distributed-fleet systems at scale.
  • Familiarity with incident/change/problem management processes (ITIL exposure helpful).
  • Knowledge of time-series anomaly detection and ML-generated signals vs rule-based alerts.
  • Experience with CI/CD tooling (GitHub Actions, Jenkins, ArgoCD, CircleCI).
  • Exposure to microservices, service mesh (Istio/Linkerd), and cloud-native patterns.
  • Experience with infrastructure-as-code (Terraform, Ansible, Pulumi).
  • Experience with fleet-scale or edge computing deployments where unattended node management matters.

What Success Looks Like at 6 Months

  • Serving on-call as primary/secondary responder, escalating with speed and context, executing runbooks independently.
  • Owning dashboards and SLOs for at least two services with measurable alert quality improvements.
  • Participating in multiple postmortems and closing assigned action items, with at least one reducing incident recurrence.
  • Automating at least three recurring manual tasks and documenting the operational toil reduction.
  • Recognized by peers as someone who closes loops and drives improvements.

Education, Hours & Type

  • Bachelor's degree in Computer Science, Engineering, or related field — or equivalent practical experience.
  • Anticipated weekly hours: 40
  • Time type: Full time

Pay & Benefits

  • The typical pay range for this role is: $72,100 - $158,620 (base annual salary). Actual offer will depend on experience, education, and geography. This position may be eligible for additional incentive programs.
  • Competitive benefits package including medical, dental, vision, paid time off, retirement options, and wellness resources. Additional benefits details provided during the hiring process.

Location & Employer

  • This role is based in Woonsocket, RI and supports distributed retail and pharmacy systems.
  • You will be employed by our confidential client and supported through the Careertakes recruiting process.

Equal Opportunity & Hiring Transparency

Careertakes and our client are Equal Opportunity Employers committed to building a diverse and inclusive workforce. We prohibit discrimination or harassment of any kind. To support a fair and efficient hiring process, AI tools may be used to assist with application review or resume screening. These tools do not replace human decision-making. Final hiring decisions are made by people.

If you have questions about how your data is used, please contact us directly.

Negotiate a higher salary! Check the salary ranges for this job type in your area.

View My Salary Range