Software Engineer – IMC Platform and Cluster Infrastructure
The Software Engineer for IMC Platform and Cluster Infrastructure at Confidential Client is responsible for Linux system administration, platform integration, and infrastructure support within semiconductor manufacturing environments. This role involves automation, deployment, and management of cluster and containerized systems, with a focus on high-performance computing, networking, and GPU-accelerated workloads. The engineer collaborates across teams and supports customer deployments, ensuring robust and efficient platform operations in mission-critical production settings.
About Careertakes
👉 Important disclosure: Careertakes is a third-party recruiting platform supporting this hiring process. If selected, you will be employed directly by our client, Semiconductor Manufacturing.
Applicants for this role may also receive access to additional matched opportunities through the Careertakes platform.
Role Overview
Confidential Client is seeking a Software Engineer for IMC Platform and Cluster Infrastructure to support Linux-based platform qualification, cluster provisioning, and deployment automation. This role sits at the intersection of systems software, infrastructure engineering, and field/platform integration. The position is based in Milpitas, CA and requires collaboration across software, hardware, and customer-facing teams.
What You’ll Do
- Administer and harden Linux platforms in enterprise and production environments, including storage, networking, and performance tuning.
- Design, implement, and maintain infrastructure automation and configuration management (Ansible, IaC, cluster provisioning).
- Develop tooling and scripts (Python, Bash) to automate installation, validation, monitoring, and operational workflows.
- Deploy, operate, and troubleshoot containerized workloads and Kubernetes clusters in production-like and customer environments.
- Work with server hardware teams to evaluate and qualify platforms (BIOS, BMC/IPMI, firmware, hardware health).
- Support GPU-accelerated compute environments (NVIDIA drivers, CUDA, performance tuning for AI/ML workloads).
- Integrate and validate high-performance networking technologies (Mellanox/NVIDIA, RDMA/RoCE, InfiniBand).
- Build and operate observability and cluster-health systems (Prometheus, Grafana, ELK, OpenTelemetry, or similar).
- Collaborate with cross-functional engineering teams, vendors, and customers to resolve mission-critical issues and perform field installations.
- Travel domestically and internationally as needed to qualify, deploy, and support platform installations.
Minimum Qualifications
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
- 2+ years of hands-on experience in Linux system administration, infrastructure/platform engineering, or deployment engineering.
- Strong scripting and software development skills (Python, Bash, or similar).
- Experience with Ansible or comparable configuration management and automation tooling.
- Familiarity with Linux networking fundamentals (TCP/IP, DNS, routing/switching) and storage/file systems.
- Experience installing, configuring, validating, and supporting software solutions in Linux-based environments.
- Proven troubleshooting and root-cause analysis skills across software, OS, networking, and hardware layers.
- Ability to work effectively across multi-functional engineering teams and to travel for customer deployments.
Preferred Qualifications
- Hands-on experience with Kubernetes administration and container orchestration.
- Experience with infrastructure-as-code, Helm, CI/CD pipelines, and automated release management.
- Background in evaluating server platforms and firmware (BIOS, BMC/IPMI).
- Knowledge of GPU ecosystems (NVIDIA drivers, CUDA) and performance tuning for compute-intensive workloads.
- Experience with high-performance networking (RDMA, RoCE, InfiniBand) and related diagnostics.
- Familiarity with observability stacks (Prometheus, Grafana, ELK) and cluster health management.
- Experience supporting air-gapped, secure, or mission-critical production deployments.
Compensation & Benefits
- Base Pay Range (Milpitas, CA): $136,300 - $231,700 annually.
- Total rewards may include performance incentives, medical/dental/vision coverage, life insurance, 401(k) with company match, employee stock purchase program, tuition reimbursement, student debt assistance, paid time off, parental leave, wellness programs, and development opportunities.
- Actual pay will depend on location, experience, skills, and qualifications. This posting complies with applicable federal and California pay transparency requirements.
Location & Travel
- Primary location: Milpitas, CA (on-site / hybrid depending on team needs).
- Willingness to travel domestically and internationally for platform qualification, customer installs, and field support is required.
Why Join (from a third-party recruiting perspective)
Careertakes partners with Confidential Client to surface this opportunity to engineers who enjoy low-level systems work, platform integration, and customer-facing technical roles. If you prefer roles that blend software development, system automation, and hands-on hardware/platform validation, this role offers high-impact work in a semiconductor manufacturing context with competitive compensation and career growth.
Equal Opportunity & Hiring Transparency
Careertakes and our client are Equal Opportunity Employers committed to building a diverse and inclusive workforce. We prohibit discrimination or harassment of any kind. To support a fair and efficient hiring process, AI tools may be used to assist with application review or resume screening. These tools do not replace human decision-making. Final hiring decisions are made by people.
If you have questions about how your data is used, please contact us directly.
Negotiate a higher salary! Check the salary ranges for this job type in your area.
View My Salary Range