Job Responsibilities
- Collaborating with stakeholders, engineers, and operational SMEs to ensure all relevant
parties are up to date with what is top of mind within the reliability service offerings - Evolve services based on customer needs and technology to ensure we remain
competitive in the market - Influence and collaborate with squads during service or platform design to proactively
prevent system failures and enhance performance - Engage with Asset/Journey squads to adopt SRE practices with a core focus to
contribute towards incident management and advocate for blameless post mortems. - Engage and influence squads with regards to observability, high availability utilising new
or existing technology and Improve disaster recovery plans. - Implement automated-based solutions to achieve high availability, efficiency, reduce
cost and performance to systems. - Coach squads on best practices within the organisation via internal forums to position
SRE fundamental knowledge and promote enterprise-wide knowledge sharing - Assist with creating and maintaining system health and performance metrics reflecting
real-time data, enabling proactive resolution and faster troubleshooting. - Collaborate and partner with DevOps engineer/coach to ensure efficient (CI/CD)
pipelines and resolve any failures or improve. - Take charge of technical leadership, engage, with squads to identify best solutions, and
support and guide Junior SRE’s. - Assist in defining and implementing metrics related to performance of services such as
SLO’s, SLI’s and SLO’s. - Defining and delivering Site Reliability Engineering technical standards in partnership
with all disciplines of software engineering. - Participate and closely work with relevant COE’s to improve release of new features to
facilitate time to market. - Ability to build and maintain strategic relationships with the business units and vendors
in order to be in sync on current ways of work and business decisions that are being
embraced - Conduct assessments within squads to measure SRE maturity, provide report and
outline a plan to assist on moving to next level with continuous feedback. - Participate and support corporate responsibility initiatives for the achievement of
business strategy. - Manage multiple concurrent objectives, projects, groups, or activities, making effective
judgements as to prioritisation and time allocation
Experience required
- Min 5 IT Experience with 3 years in relevant technology or domain
- Working Experience of Operating System (Linux or Windows)
- Knowledgeable with microservices and containerization; K8s or Docker
- Troubleshooting and rout cause Analysis
- SRE Best practices
- In-depth knowledge of DevOps framework
- Experience and knowledge of programming languages(C#, Java, Python, Bash)
- Proactivity in seeking Improvement opportunities
- Experience with troubleshooting production systems/applications
Qualifications required
- Degree or Diploma in IT
