Careers · Engineering · Taipei

Site Reliability Engineer

Our uptime number isn't a marketing line — it's a promise to customers whose logins and payments depend on us. We're looking for an SRE who treats 99.9% as an engineering problem: designed for, measured, and defended, not hoped for.

Location
Taipei
Team
Engineering — Infrastructure
Type
Full-time
Workplace
On-site (hybrid)

About the role

Reliability is a feature.

You'll own the health of the platform — the infrastructure that messaging, voice, email, and the AI agents all run on — across multiple regions.

That means observability people actually use, failover that works when a zone has a bad day, capacity that's there before a customer's launch, and incident practices that turn a 2 a.m. page into a fix that sticks. You'll partner with the product engineering teams to make reliability part of how things are built, not a cleanup crew after the fact.

What you'll do

  • Own observability end to end — metrics, logs, tracing, and dashboards that make the system's state obvious.
  • Design and test multi-region failover and disaster-recovery so a single zone's outage doesn't reach customers.
  • Set and track SLOs with the product teams, and hold the line on error budgets.
  • Lead incident response and blameless postmortems; drive the follow-ups that prevent repeats.
  • Improve deployment safety — progressive rollouts, automated rollback, and infrastructure as code.
  • Plan capacity ahead of customer launches and seasonal peaks.

What we're looking for

  • 4+ years in SRE, production, or platform engineering for systems that can't go down.
  • Solid with containers and orchestration (Kubernetes), infrastructure as code (Terraform or similar), and CI/CD.
  • Fluent in an observability stack (Prometheus, Grafana, OpenTelemetry, or equivalents) and comfortable scripting/automating in Go or Python.
  • You've run real incidents and written the postmortem that made things better.
  • Calm under pressure, and biased toward fixing the cause rather than the symptom.
  • Working proficiency in English.

Nice to have

  • Experience operating multi-region, latency-sensitive systems.
  • Security or edge-protection experience (WAF, rate limiting, DDoS mitigation).
  • A background that includes writing the services you now operate.

Ready to apply?

Tell us about an incident you turned into a fix.

Résumé and a short note. A real engineer reads every application.

Apply

Submit your application

Fill in the form and add a link to your résumé or portfolio. A real person on our team reviews every application — usually within a few days.

✓ We reply to every applicant✓ No account or login required✓ Your details go straight to our hiring team