careertakescareertakes

Software Engineer, Spark Platform

Confidential ClientTransportation
Entry LevelFull-timeHybridN/A
San Francisco, California
15 Days Ago

The Software Engineer on the Spark Platform team at Confidential Client is responsible for building and operating a large-scale in-house Apache Spark platform that supports data, analytics, and ML workloads company-wide. The role involves optimizing Spark runtime, managing multi-tenant scheduling and cluster lifecycle automation on Kubernetes, and developing observability and incident automation tools to ensure platform reliability and scalability. This hybrid position requires collaboration across teams and offers opportunities for deep technical ownership in distributed systems and platform operations.

Boost your chances. Upload your resume to see your match score.

Unlock Match Rate
Description

About Careertakes

👉 Important disclosure: Careertakes is a third-party recruiting platform supporting this hiring process. If selected, you will be employed directly by our client, Software Development.

Applicants for this role may also receive access to additional matched opportunities through the Careertakes platform.


What You’ll Do

As a Spark Platform Engineer on this team you will design, build, and operate an in‑house Apache Spark platform that runs production workloads at scale. Day‑to‑day responsibilities include:

  • Build and operate an in‑house Spark platform spanning runtime, scheduler, reliability, and user-facing tooling.
  • Lead runtime upgrades, performance tuning, and shuffle/runtime reliability work for Spark at scale.
  • Drive multi‑tenant scheduling, executor bin‑packing, and cost‑aware placement to efficiently serve many consumer teams.
  • Implement and own cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling.
  • Build observability, SLO/SLI definitions, and incident automation so the platform is debuggable end‑to‑end and on‑call load is sustainable.
  • Work across system layers — scheduler, runtime, orchestration, and user tooling — and partner closely with platform consumers.
  • Prioritize incremental rollouts, measurement, and reducing operational toil.

What We’re Looking For

The ideal candidate will bring hands-on experience operating distributed data platforms and strong engineering judgment.

  • B.S., M.S., or PhD in Computer Science or equivalent experience.
  • 2+ years operating production distributed systems (batch/big‑data platforms).
  • Experience operating Apache Spark at scale (AWS EMR, Databricks, or in‑house Spark deployment) with emphasis on platform operations (runtime upgrades, cluster lifecycle, shuffle, observability, multi‑tenant scheduling).
  • Hands‑on experience with Kubernetes in production — controllers, operators, custom resources, and multi‑tenant failure modes.
  • Familiarity with batch or big‑data schedulers (YuniKorn, Volcano, Kueue, or similar) and/or Spark-on-Kubernetes operator.
  • Experience with observability stacks (Prometheus, OpenTelemetry, distributed tracing, structured logging) and defining SLOs/SLIs.
  • Comfortable in cloud environments (AWS preferred): VPC networking, instance lifecycle, spot/preemptible markets, autoscaling.
  • Proficiency in one or more of: Python, Go, Scala, Java; strong SQL fluency.
  • Bias toward measurement-driven changes, incremental rollouts, and reducing operational toil.
  • Located in or willing to relocate to San Francisco, CA for this hybrid role.

Compensation & Benefits

The position is full time. The national base pay ranges provided for this role in the United States are listed below. Final salary is determined by job-related factors including skills, experience, work location, and market conditions. Base salary is localized to the employee’s work location.

  • I4: $130,600 — $192,000 USD (base)
  • I5: $159,800 — $235,000 USD (base)
  • I6: $193,800 — $285,000 USD (base)

In addition to base pay, compensation may include opportunities for equity grants. Our client provides a comprehensive benefits package, which may include: 401(k) with employer matching, parental leave, medical/dental/vision coverage, paid time off and paid sick leave in compliance with applicable law, disability and basic life insurance, commuter benefits match, and mental health resources. Details and eligibility depend on role and location.


Work Location & Schedule

This is a hybrid role based in San Francisco, CA. Candidates must be located in or willing to relocate to San Francisco for the role.


Why Apply via Careertakes

Careertakes is a third‑party recruiting partner for this hiring process. Applying through Careertakes gives you the chance to be matched to this role and to other opportunities we screen for on behalf of our clients. Our team follows third‑party recruiting best practices and works to present your profile appropriately to the hiring team.


Inclusive Hiring & Accommodations

Careertakes and our client are committed to building an inclusive workforce. We consider applicants with arrest or conviction records in a manner consistent with applicable laws. If you need an accommodation during the recruiting process, please let your recruiting contact know.


Equal Opportunity & Hiring Transparency

Careertakes and our client are Equal Opportunity Employers committed to building a diverse and inclusive workforce. We prohibit discrimination or harassment of any kind. To support a fair and efficient hiring process, AI tools may be used to assist with application review or resume screening. These tools do not replace human decision-making. Final hiring decisions are made by people.

If you have questions about how your data is used, please contact us directly.

Negotiate a higher salary! Check the salary ranges for this job type in your area.

View My Salary Range