Software Engineer, SRE

Bet3651

ManchesterOn-siteEst. £64k - £85k (similar roles)full timePlatform EngineeringPosted 1w agoEarly applicant likely

This role is aggregated from the employer's public careers feed - the button opens their application page.

As a Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing for, and progressing the performance and availability of our critical systems. Using your engineering expertise, you will implement solutions that enhance reliability, including service instrumentation with OpenTelemetry and improved logging practices. You will leverage AI tools and LLM platforms in your daily work to reduce toil, drive autonomous operations, and optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role is eligible for inclusion in the Company's hybrid working from home policy. - Excellent knowledge of programming languages including Python, Golang and JavaScript. - Knowledge and experience of modern software development techniques and lifecycles. - Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. - Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty. - Proficiency in shell scripting for automation and system management tasks. - Experience with Infrastructure as Code (IaC), automation and orchestration tools such as Ansible and Terraform. - Prior experience working in a large scale, 24/7 enterprise where system uptime and stability is of paramount importance to the business. - An AI-native engineering approach, with hands-on experience using LLM platforms and coding assistants to improve productivity and quality, and the ability to integrate AI-driven telemetry for advanced observability, predictive insights and root-cause analysis.

About Bet3651

Bet3651 builds facial recognition and identity verification software - specifically their Cerberus platform which handles identity, compliance and operational functions in regulated sectors. The company also maintains HVAC systems across a property portfolio.

All Bet3651 jobs

Stop searching - get matched.

Upload your CV and our AI will surface roles like this one, with the reasons they fit you.

Upload CV - it's free