Core platform Engineer /SRE
Enormous Enterprise LLC
Core platform Engineer /SRE
Sunnyvale, CA - Locals only
Long Term Contracct
Strong hands-on experience in at least one of these areas:
- GPU-based hardware (management, troubleshooting, at-scale operations)
- SDN (OVN, OVS)
- Storage (Lightbits, VAST, Pure Storage
What You'll Be Working On:
As a CORE PE, you will engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones.
Each morning starts with a structured review of overnight alerts and system performance metrics - identifying any anomalies, triaging what needs attention.
You will collaborate with your team in a morning stand-up meeting to discuss ongoing projects, recent incidents, and priorities for the day
Your tasks will include automating routine processes, analyzing system logs, and developing tools to enhance our monitoring capabilities.
You'll spend part of your day working closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment.
Throughout the day, your focus is on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers.
By day's end, you will document your work, share insights with your team, and plan for the next day's challenges, always with a customer-centric mindset.
What Youll Bring to the Team:
Strong experience with architecture, design patterns, reliability and scaling of new and current systems.
Experience leading and commanding incidents, including driving root cause analysis, coordinating cross-functional teams, and ensuring follow-through on corrective actions.
Experience building observability from the ground up defining SLOs/SLIs, closing monitoring gaps, and implementing alerting strategies that catch failures before customers do.
Proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystem.
Experience writing high quality code with at least one programming language (Python, Go, or similar).
Experience with system-level debugging, including kdump, and kernel panic analysis
Proficiency in Infrastructure as Code tooling (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure
Experience with TCP/IP and network programming.
Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms
Hardware and GPU troubleshooting experience (nice to have).
Exposure to OVN/OVS-based networking stack (nice to have).
Strong communication skills.
Sunnyvale, CA - Locals only
Long Term Contracct
Strong hands-on experience in at least one of these areas:
- GPU-based hardware (management, troubleshooting, at-scale operations)
- SDN (OVN, OVS)
- Storage (Lightbits, VAST, Pure Storage
What You'll Be Working On:
As a CORE PE, you will engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones.
Each morning starts with a structured review of overnight alerts and system performance metrics - identifying any anomalies, triaging what needs attention.
You will collaborate with your team in a morning stand-up meeting to discuss ongoing projects, recent incidents, and priorities for the day
Your tasks will include automating routine processes, analyzing system logs, and developing tools to enhance our monitoring capabilities.
You'll spend part of your day working closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment.
Throughout the day, your focus is on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers.
By day's end, you will document your work, share insights with your team, and plan for the next day's challenges, always with a customer-centric mindset.
What Youll Bring to the Team:
Strong experience with architecture, design patterns, reliability and scaling of new and current systems.
Experience leading and commanding incidents, including driving root cause analysis, coordinating cross-functional teams, and ensuring follow-through on corrective actions.
Experience building observability from the ground up defining SLOs/SLIs, closing monitoring gaps, and implementing alerting strategies that catch failures before customers do.
Proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystem.
Experience writing high quality code with at least one programming language (Python, Go, or similar).
Experience with system-level debugging, including kdump, and kernel panic analysis
Proficiency in Infrastructure as Code tooling (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure
Experience with TCP/IP and network programming.
Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms
Hardware and GPU troubleshooting experience (nice to have).
Exposure to OVN/OVS-based networking stack (nice to have).
Strong communication skills.
Reference: 3144673356