Fingerprint
Senior Site Reliability Engineer
Fingerprint
$152k - $205k
Worldwide (Remote)
AWS
Terraform
Go

Senior Site Reliability Engineer

Overview

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence.

Job Description

Fingerprint is a globally dispersed, 100% remote company. We were named on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies.

Responsibilities

  • - Own the reliability of core production systems end to end — you instrument them, set targets for them, operate them, and are accountable for how they behave under real traffic.
  • - Define and maintain SLIs and SLOs for the critical paths you own, wire them into dashboards and alerts, and use error budget burn as the evidence base for what gets fixed next.
  • - Drive alert quality: raise signal, kill noise, and close the gap where customers notice a problem before our monitoring does.
  • - Take a lead role in incident response — investigate systematically across service boundaries, restore service, and write postmortems that produce follow-ups people actually complete.
  • - Build secure, resilient, and cost-efficient infrastructure, with explicit attention to failure modes: timeouts and retries, backpressure and load shedding, graceful degradation, and blast radius containment.
  • - Do capacity and performance work with real data — load testing, profiling, saturation analysis, and headroom planning ahead of growth rather than after an incident.
  • - Improve change safety: progressive delivery, automated rollback, meaningful pre-production signal, and deployment practices that make shipping boring.
  • - Manage infrastructure through code and configuration (we primarily use Terraform), consistently applying patterns that align with our overall service architecture.
  • - Design, write, and ship software and developer-facing tooling that reduces toil and makes operating services straightforward for the engineers who own them.
  • - Run deliberate failure testing — game days and chaos exercises, staging first — to find the gaps and safe limits before customers do.
  • - Partner with product engineering teams on production readiness for new and high-risk services: capacity, failure modes, rollback plans, runbooks, and on-call handoff.
  • - Teach through review rather than gatekeeping.
  • - Participate in the on-call rotation, and improve it: better runbooks, clearer escalation, less pager fatigue for everyone in it.
  • - Approach all engineering work with a security lens — actively looking for vulnerabilities in your own work and in peer reviews.
  • - Act as the go-to person for hard production problems in your area, and mentor engineers through code review, pairing, and design feedback so operational knowledge doesn't silo.

Required Skills

  • - 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering within primarily cloud-based environments (AWS preferred), with meaningful time spent responsible for systems in production.
  • - A track record of owning a system end to end — you've designed something significant, shipped it, operated it, and lived with the consequences when it misbehaved.
  • - Hands-on experience defining and operating against SLIs, SLOs, and error budgets — not just reading the book, but getting targets adopted and acted on.
  • - Strong incident skills: you've led or been a primary responder on high-severity, customer-facing incidents, and you've improved how an organization learns from them.
  • - Depth in distributed systems failure modes in high-throughput, low-latency environments — cache and database saturation, cascading failure, retry storms, capacity limits, degradation and load shedding.
  • - Depth in cloud infrastructure fundamentals: networking, load balancing, containerization (EKS/Kubernetes), and distributed systems.
  • - Strong hands-on experience managing infrastructure through code and configuration (Terraform or equivalent).
  • - Solid programming skills in Go, Python, or a comparable language — you write real, production-ready software and can ship the fix rather than only recommend it.
  • - Fluency with observability tooling (Datadog, Prometheus, Grafana, OpenTelemetry, or similar), including instrumenting systems yourself rather than inheriting dashboards.
  • - Hands-on experience operating Redis/ElastiCache in production — including cluster/shard management, failover behavior, memory eviction policies, and scaling strategies.
  • - Fluency with software engineering best practices: source control, code review, comprehensive test coverage across edge cases and errors, and safe deployment.
  • - A high level of personal ownership and autonomy, with real experience working without clearly defined requirements.
  • - Pragmatism over purity — you know reliability competes with delivery, can make the case for the right investment at the right time, and can say when a risk is acceptable.
  • - Strong written and verbal communication in English — clear technical design docs, PR reviews, incident updates, and postmortems that bring engineers outside your team along on a decision.
  • - AI-native by default. You use AI tools as a normal part of how you investigate incidents, analyze telemetry, write runbooks, and build tooling — and you have opinions, from experience, about where they help and where they don't yet.

About the company

Identify every visitor. Stop fraud, detect bots, or delight customers. Identify good and bad visitors with industry-leading accuracy - even if they're anonymous.


All Job Openings at Fingerprint