Lead Site Reliability & Security Enginee
About the Role
Are you a seasoned SRE and Security expert ready to make a significant impact? Join a well-funded, Series A/B SportsTech innovator building a cutting-edge, commission-free, peer-to-peer sports prediction exchange. We're seeking a **Lead Site Reliability & Security Engineer** to anchor our hands-on, cross-functional engineering team in New York. This isn't just a role; it's a mission to secure and scale a regulated, exchange-grade trading platform. You'll take full ownership of disaster recovery, pioneering observability frameworks, and embedding security-by-design at every layer. Your expertise will ensure our high-throughput platform remains resilient, compliant, and always available for our users.
What You'll Do
**Architect & Own Core Systems:** Design, implement, and lead the charge on disaster recovery, reliability, and cutting-edge observability frameworks crucial for our regulated trading platform.
**Champion Security-by-Design:** Spearhead and implement proactive security infrastructure initiatives, ensuring a robust security posture is intrinsically woven into every facet of our engineering ecosystem.
**Drive Automation & Efficiency:** Revolutionize our recovery and reliability processes, leveraging sophisticated infrastructure-as-code to build self-healing and auto-scaling systems.
**Cross-Functional Impact:** Collaborate closely with product and compliance teams, translating regulatory requirements into resilient technical solutions that guarantee exchange-grade uptime and impeccable auditability.
**Bridge SRE & Security:** Serve as the critical link, fostering seamless integration and communication between reliability engineering and security operations to create a unified and impenetrable platform.
What We're Looking For
**Dealbreakers (must-haves):**
**5+ Years of Blended Expertise:** A minimum of 5 years of hands-on experience combining Site Reliability Engineering (SRE) and security engineering, specifically within dynamic cloud environments.
**Regulated Platform DR Mastery:** A verifiable track record of designing, implementing, and managing robust disaster recovery and business continuity plans for **regulated trading platforms**.
**Required skills & experience:**
**Leadership & Collaborative Influence:** Demonstrated leadership in guiding technical initiatives, mentoring engineers, and fostering effective cross-functional collaboration, with exceptional clarity in communicating complex concepts to both technical and non-technical stakeholders.
**Observability Visionary:** Extensive experience in architecting and deploying comprehensive observability strategies (metrics, logs, tracing) that empower proactive monitoring and ensure reliability at enterprise scale.
**IaC & Cloud Orchestration:** Expert-level proficiency with Infrastructure as Code tools, specifically **Terraform** and **AWS CloudFormation**, for managing complex cloud infrastructure.
**Incident Response & Resilience:** Proven strength in incident response, including swift root-cause analysis, effective mitigation, and meticulous post-incident reporting.
**Security-by-Design Evangelist:** Deep-seated knowledge of security best practices spanning all infrastructure layers, coupled with an unwavering security-by-design mindset.
**Event-Streaming Expertise:** Hands-on experience with high-throughput event-streaming systems, particularly **Kafka** or comparable technologies.
**Database Fluency:** Practical, hands-on experience managing and optimizing both relational databases (**PostgreSQL**) and NoSQL databases (**MongoDB**).
Compensation & Benefits
**Salary:** $185,000 – $235,000 USD annually, commensurate with experience.
Visa sponsorship is **not available** for this role.
Location
This role is based in **New York, NY**. On-site or hybrid presence is expected. This is not a fully remote position.