Descrição da vaga
<p><strong>What you’ll be doing</strong></p>
<ul>
<li>Work with a team of DevOps and DBA professionals</li>
<li>Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future</li>
<li>Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices</li>
<li>Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)</li>
<li>Own weekday on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews</li>
<li>Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding</li>
<li>Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation</li>
<li>Take ownership and responsibility for our cloud operation activities</li>
<li>Liaise with external security agencies for annual audits as well as perform our own internal security sweeps</li>
<li>Aid in reconfiguring existing architecture to allow for rapid deployments to new countries</li>
<li>Mentoring less experienced team members</li>
</ul>
<p><strong>What you’ll bring</strong></p>
<ul>
<li>3+ years DevOps / SRE / platform engineering experience</li>
<li>Must be based in Europe </li>
<li>Experience independently leading the planning and deployment of a project</li>
<li>Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production</li>
<li>Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued</li>
<li>Experience with Infrastructure-as-Code, particularly Terraform</li>
<li>Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus</li>
<li>Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry</li>
<li>Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus</li>
<li>Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions</li>
<li>Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns</li>
<li>Experience defining SLIs and SLOs and using them to inform reliability work</li>
<li>Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments</li>
<li>Solid networking knowledge, especially the TCP / IP stack and HTTP protocol</li>
<li>Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments</li>
<li>A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached</li>
<li>Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous</li>
</ul>
<p><strong>Our stack</strong></p>
<ul>
<li>Languages: Java / Spring Boot, Node.js, Python, JavaScript</li>
<li>Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community</li>
<li>Cache: ElastiCache, Redis, Valkey</li>
<li>Messaging: Apache RocketMQ, AutoMQ, Kafka</li>
<li>Networking & Proxy: Nginx, Kong, Cilium, eBPF</li>
<li>Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm</li>
<li>Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3</li>
<li>CI/CD: Jenkins, GitHub Actions</li>
<li>Metrics: Prometheus, Mimir, Grafana, Alertmanager</li>
<li>Logs: Loki, Vector</li>
<li>Traces: Tempo, OpenTelemetry, Alloy</li>
<li>Profiling: Pyroscope</li>
<li>RUM: Grafana Faro, OpenTelemetry SDK</li>
<li>Infrastructure as Code: Terraform</li>
<li>CDN & Edge: Cloudflare, AWS CloudFront</li>
<li>AWS CloudWatch</li>
</ul>
<p><strong>What’s in it for you</strong></p>
<ul>
<li>Sporty is a remote first company in pursuit of sustainability</li>
<li>A competitive salary + individual performance based bonuses every quarter</li>
<li>28 days paid annual leave</li>
<li>Our core working hours are 10am-3pm in your local time zone with flexibility outside of this</li>
<li>Referral bonuses & flash bonuses</li>
<li>Top of the line equipment</li>
<li>Annual company retreats to provide great internal networking opportunities</li>
</ul>
<p><strong>Interview process</strong></p>
<ul>
<li>Remote video screening with our Talent Acquisition Team</li>
<li>Online assessment via Hackerrank</li>
<li>Remote video interview with 3 x Team Members (45 mins each, not separate days)</li>
</ul>
<p>If you’re interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.</p>