Remote Senior Site Reliability Engineer (US or Canada)
Senior Site Reliability Engineer
Remote — United States or Canada
$160-210K base + equity
About the Company
We’re partnering with a well-funded Series B AI company building an enterprise platform at the intersection of voice AI, large language models, and real-time customer interactions. The company has raised more than $100M, supports large enterprise customers, and has been running AI agents in production for years—not simply adding AI to an existing product.
The engineering organization is intentionally lean and senior. This is a 100% remote company across the U.S. and Canada, with engineers owning multiple initiatives and operating with significant autonomy.
The Role
This is not a traditional infrastructure-only SRE position.
You’ll join a small, senior platform/SRE team responsible for the systems that enable the broader engineering organization to build, test, deploy, observe, and operate an AI-native production platform.
The team owns areas including:
- CI/CD and developer experience
- Platform engineering and reliability
- Developer tooling and paved paths
- Observability and incident management
- Kubernetes and cloud infrastructure
- Cloud cost visibility and optimization
- AI-agent harness engineering
- Sandboxes, guardrails, validation, and agent-first development workflows
The hiring manager is particularly interested in engineers with strong DevEx, CI/CD, platform engineering, or SRE backgrounds rather than candidates whose experience is primarily compute, networking, data centers, or traditional infrastructure administration.
What You’ll Do
- Own and evolve CI/CD systems end to end, including architecture, caching, build/test performance, deployment workflows, and developer self-service.
- Build tooling and platform capabilities that reduce engineering toil and help developers ship safely and quickly.
- Help extend an internal AI-agent harness used to support autonomous development workflows, including CI integration, sandboxing, guardrails, and validation.
- Improve reliability and operability for a high-volume, real-time AI platform.
- Build and maintain infrastructure using Terraform, Kubernetes, and Helm.
- Improve observability across logs, metrics, traces, monitoring, and alerting.
- Partner directly with software engineers to troubleshoot production and development issues, including making changes within application code where needed.
- Help establish patterns and technical standards across the engineering organization.
- Participate in an on-call rotation focused on base infrastructure; after-hours pages are uncommon.
What We’re Looking For
- 6+ years of experience in SRE, platform engineering, DevEx, developer infrastructure, or related software-development enablement roles.
- Strong experience owning CI/CD platforms end to end, rather than simply maintaining existing pipelines.
- Experience with CI/CD concepts such as caching, architecture, build/test optimization, deployment ergonomics, and developer self-service.
- Strong enough software engineering fundamentals to succeed in a coding-oriented technical interview.
- Working familiarity with:
- TypeScript / Node.js
- Python
- Terraform
- Kubernetes / Helm
- Practical observability experience across logs, metrics, tracing, monitoring, alerting, and incident management.
- Experience working successfully on distributed or fully remote engineering teams.
- Understanding of how modern LLMs and AI development tools work.
- Hands-on use of tools such as Claude, Cursor, Copilot, or similar—and the judgment to know when AI-generated output should not be trusted without deeper evaluation.
The environment currently includes TypeScript, Node.js, Python, Kubernetes, Helm, GCP, GitLab CI, Terraform, Datadog, Prometheus, and Grafana.
Especially Interesting Backgrounds
We’d be particularly interested in engineers who have worked on:
- Developer platforms or internal developer infrastructure
- CI/CD infrastructure at meaningful scale
- Developer productivity / DevEx
- Cloud-native SaaS platforms
- AI or ML infrastructure
- Platforms for autonomous or semi-autonomous AI agents
- Large-scale GCP environments
- Real-time communications, telephony, SIP, or FreeSWITCH
Experience in voice AI or conversational AI is helpful, but not required.
What This Role Is Not
This is probably not the right fit if your background is primarily:
- Network engineering
- Data center or on-prem infrastructure
- Cloud compute administration without meaningful DevEx or software engineering ownership
- Infrastructure operations with little hands-on coding
- Maintaining CI/CD pipelines without having designed or owned the platform behind them
Team & Culture
You’ll join a small, senior engineering organization with staff- and principal-level engineers and substantial individual ownership. There are no junior engineers on the team.
The culture rewards people who:
- Take ownership and drive projects forward.
- Are comfortable operating across multiple technical domains.
- Will tackle both large architectural problems and unglamorous operational work.
- Move quickly, experiment, learn from mistakes, and iterate.
- Can critically evaluate their own decisions and respond well to feedback.
Compensation & Location
Base salary: approximately $160,000-210,000 USD, depending on level and location, plus competitive equity. Compensation may vary for Canadian employees.
Location: Fully remote within the United States or Canada, working primarily across U.S. time zones.
