Scalable GmbH
(Senior) Cloud Site Reliability Engineer (Scalability) (m/f/x)
- Location
- München, BY, Germany
- Work model
- Hybrid
- Seniority
- Senior
- Employment
- FullTime
- Posted
- Added to Codestelle
Language requirements
- German
- Not specified
- English alone
- Not specified
Based on explicit wording in the listing. “Not specified” does not mean a language is optional.
About this role
Our team's mission is to ensure that teams can scale up their services while being performant, reliable and cost optimised through self-service tooling and transferring best-practices for observability, load-testing, services discoverability, chaos engineering and FinOps.
Scalable Capital was built in the cloud from day one. Our services currently run on various AWS services like ECS, Fargate and Lambda and are distributed across multiple accounts. We embrace a DevOps culture where the development teams manage their CI/CD pipelines and cloud infrastructure for their services themself. Our Scalability Engineering Team focuses on providing everything necessary to ensure Reliability and Scalability of our services and storages while being cost efficient. This includes solutions for Monitoring, AutoScaling, Load Testing, Chaos Engineering and FinOps.
- Shape the way how Scalable runs microservices in the most performant, secure and cost efficient way.
- Collaborate with cross-functional teams to identify and understand scalability requirements for our platform, both in terms of user growth and increasing data volume.
- Design and rollout Monitoring best practices in Datadog including SLI, SLO and SLAs
- Research and develop service and storage improvements by using serverless technologies and optimise our services and CICD to optimise scalability, cost and performance
- Develop and maintain internal tooling around Monitoring, Developer Portal and Load Testing
- Mentor and enable our software development teams to further foster our DevOps culture by educating them and providing reusable and unified building blocks which can be used to improve scalability, reliability and performance of existing and new services
- Stay up-to-date with the latest industry trends, tools, and techniques related to scalability and performance engineering.
- Design and implement best practices around auto scaling of our infrastructure
- Run chaos engineering experiments to improve resilience of our services
- Multiple years of experience with AWS and infrastructure as code (preferably Terraform)
- Solid experience in Monitoring, Container Orchestration and Microservice setups is required
- Good working knowledge with Python, at least one additional general purpose programming language (Preferably Java/Kotlin or JavaScript/Node.js) and build automations tools
- Solid understanding of scalable system design principles, distributed systems, and cloud technologies
- Experience with GitHub Actions and Jenkins would be beneficial
- A passion for automating, improving processes and working together with other developer teams
- A degree in a relevant field of study (e.g. computer science, engineering, sciences) or work experience in a role that typically requires a university degree
- Full professional proficiency in English and the ability to communicate concisely in an international English-speaking environment
- Excellent communication and collaboration skills, with the ability to work effectively in a cross-functional team environment.