Remote job
Remote confirmedSenior Software Engineer, Infrastructure
Stream
Our assessment
- Our reading of the full posting text confirms it: fully remote.
- 2 more open roles from this employer in our index. 1 of them fully remote.
This section only: calculated automatically by nomado24, from our own job index and our own reading of the posting text. Not stated by the employer.
Job description
SENIOR SOFTWARE ENGINEER, INFRASTRUCTURE
THE ROLE
We are hiring a Senior Software Engineer to help rebuild the platform underneath Stream. Over the next year the infrastructure team is moving from AWS to GCP, moving onto Kubernetes, and relocating 35 to 40 Postgres shards off managed RDS to self-hosted, while the platform keeps serving billions of API requests a month. You will own parts of that outright.
This is a small, senior team without the support structures of a large organisation. You will write code most of the time and make infrastructure calls on your own. Success looks like systems that scale predictably under load, cloud spend that falls per unit of traffic, and migrations that land without incident.
This is a full-time job opening based in Amsterdam (3 days hybrid) or remote in Europe/EU.
ABOUT STREAM
Stream powers real-time Chat https://getstream.io/chat/, Video https://getstream.io/video/, Activity Feeds https://getstream.io/activity-feeds/, and AI Moderation https://getstream.io/moderation/ for billions of end-users across thousands of apps, from Strava and Bumble to eBay and Patreon. Our platform processes billions of API requests per month and supports applications with millions of concurrent users, while delivering highly reliable, low-latency services and a great developer experience.
WHAT YOU WILL DO
- Design, build and operate infrastructure for real-time systems carrying millions of concurrent connections and billions of monthly API requests.
- Drive Kubernetes end to end: cluster architecture, workload design and the migration of existing services. You will be designing clusters, not operating someone else's.
- Re-architect workloads as part of the AWS to GCP migration, for cost and performance rather than a lift and shift.
- Own cloud cost and efficiency work: find the levers, measure them against real spend and utilisation data, and show what moved.
- Write production Go and Python: internal services, platform tooling and automation that change how product and SDK engineers deploy, observe and debug.
- Lead post-migration tuning and capacity planning, closing the loop between the architecture you chose and what production actually does.
- Work with backend, video and moderation engineers on system design, reliability targets and tradeoffs that cross service boundaries.
- Take part in on-call, incident response and root cause analysis, and turn what you find into durable fixes.
WHAT WE ARE LOOKING FOR
- 5+ years in infrastructure, platform, DevOps or SRE engineering, with clear depth in infrastructure over application development.
- A software engineering background. You have built systems, not only configured them. Production coding experience in Go or Python. Scripting-only backgrounds are not a fit.
- Kubernetes at meaningful production scale, past operations: you have driven cluster strategy, designed workloads, or led a migration, and you have tuned what came out the other side for cost and efficiency.
- Cloud cost or efficiency optimisation you personally led on AWS or GCP, with an outcome you can put a number on. FinOps practice is a plus.
- Direct experience running high-scale, high-load production systems.
- Strong cloud fundamentals across networking, compute, storage and IAM, and the habit of asking why a system behaves the way it does instead of accepting the default.
- Comfortable in a small team: leading a project and reviewing a PR in the same week.
- AI tooling already in your engineering workflow. Applied use, not familiarity.
BONUS POINTS
- Both AWS and GCP, and migration experience between providers.
- PostgreSQL at scale: sharding, replication strategy, partitioning tradeoffs, ideally self-hosted.
- Real-time systems: WebSockets, WebRTC, streaming or other persistent-connection workloads.
- The wider stack: CockroachDB, Redis, Terraform, and a Prometheus-based observability stack.
- An API-first or infrastructure company at scaleup stage.
- Open source contributions to infrastructure or platform tooling.
- Writing or talks on cloud, platform or distributed systems.
- Formal FinOps practice, or owning cloud commitment and reservation strategy.
- Work on developer-facing API or SDK products.
OUR STACK
- Go, gRPC, RocksDB, Python
- PostgreSQL, RabbitMQ
- GCP
- Grafana, Prometheus, ELK (Elasticsearch and Kibana)
- Jaeger and Tempo for distributed tracing, Datadog
- Redis, Memcached
- Claude Code, Cursor
YOU WILL THRIVE HERE IF
- You want infrastructure problems at a scale most engineers never touch, and the autonomy to own them.
- You ship fast and learn fast, including when it is hectic.
- You are self-directed and comfortable working with a globally distributed team across time zones.
YOU PROBABLY WILL NOT IF
- You want tightly scoped tickets and step-by-step direction.
- You need a calm, highly predictable environment.
- You would rather wait for a defined …
This role is provided by an external source. Applications are handled on the source website.
