Site Reliability Engineer

  • Full Time
  • Sandton
  • Posted 1 year ago

Job Responsibilities

  • Collaborating with stakeholders, engineers, and operational SMEs to ensure all relevant
    parties are up to date with what is top of mind within the reliability service offerings
  •  Evolve services based on customer needs and technology to ensure we remain
    competitive in the market
  • Influence and collaborate with squads during service or platform design to proactively
    prevent system failures and enhance performance
  • Engage with Asset/Journey squads to adopt SRE practices with a core focus to
    contribute towards incident management and advocate for blameless post mortems.
  •  Engage and influence squads with regards to observability, high availability utilising new
    or existing technology and Improve disaster recovery plans.
  • Implement automated-based solutions to achieve high availability, efficiency, reduce
    cost and performance to systems.
  •  Coach squads on best practices within the organisation via internal forums to position
    SRE fundamental knowledge and promote enterprise-wide knowledge sharing
  • Assist with creating and maintaining system health and performance metrics reflecting
    real-time data, enabling proactive resolution and faster troubleshooting.
  •  Collaborate and partner with DevOps engineer/coach to ensure efficient (CI/CD)
    pipelines and resolve any failures or improve.
  •  Take charge of technical leadership, engage, with squads to identify best solutions, and
    support and guide Junior SRE’s.
  • Assist in defining and implementing metrics related to performance of services such as
    SLO’s, SLI’s and SLO’s.
  • Defining and delivering Site Reliability Engineering technical standards in partnership
    with all disciplines of software engineering.
  • Participate and closely work with relevant COE’s to improve release of new features to
    facilitate time to market.
  •  Ability to build and maintain strategic relationships with the business units and vendors
    in order to be in sync on current ways of work and business decisions that are being
    embraced
  • Conduct assessments within squads to measure SRE maturity, provide report and
    outline a plan to assist on moving to next level with continuous feedback.
  • Participate and support corporate responsibility initiatives for the achievement of
    business strategy.
  • Manage multiple concurrent objectives, projects, groups, or activities, making effective
    judgements as to prioritisation and time allocation

Experience required 

  • Min 5 IT Experience with 3 years in relevant technology or domain
  •  Working Experience of Operating System (Linux or Windows)
  • Knowledgeable with microservices and containerization; K8s or Docker
  • Troubleshooting and rout cause Analysis
  • SRE Best practices
  • In-depth knowledge of DevOps framework
  • Experience and knowledge of programming languages(C#, Java, Python, Bash)
  • Proactivity in seeking Improvement opportunities
  • Experience with troubleshooting production systems/applications

Qualifications required

  • Degree or Diploma in IT

Disclaimer

By clicking Send application you confirm the following:

Personal Information

1. That you have no objection to us retaining your personal information in our database for future matching.

Other Suitable Opportunities

2. Should suitable opportunities arise we will contact you and request your consent to submit your CV to a specific client for a specific purpose.

Correct Information

3. That the information you have provided to us is true, correct and up to date.

Upload your CV/resume or any other relevant file. Max. file size: 32 MB.