Site Reliability Engineer
4 days ago
Remote job, Srbija
NOVACARD
Рад на даљину
Пуно радно време
200.000 MX$ Уговор
Бесплатно путем е-поште или Google
Сачувајте овај посао и одржавајте претрагу организованом
Направите бесплатан налог да бисте сачували послове, креирали упозорења и вратили се на овај унос са своје контролне табле.
Бесплатно путем е-поште или Google
*At
* *[NOVACARD](https://novacard.mx/)**, we’re redefining how people use credit. We are the
* *first interest-free and no-annual-fee credit card in Mexico, designed to simplify personal finances and give users complete control
- all from a mobile app. With NOVACARD, users can access up to
* *$200,000 MXN in credit**, only pay when they use it, and manage everything digitally in under 5 minutes. Our mission is to empower people to make smarter financial decisions by offering flexibility, transparency, and the freedom they need to reach their goals. Simple finances, big goals.*
About the Role
*
* We’re looking for a Site Reliability Engineer (SRE) to ensure the stability, performance, and reliability of our critical production systems. You’ll work at the intersection of development and operations — building automation tools, improving observability, and preventing incidents before they occur.
Key Responsibilities
- Ensure the stability, performance, and fault tolerance of production systems.
- Develop and maintain infrastructure automation and observability tools.
- Monitor system health, respond to incidents, and perform root cause analysis (RCA).
- Collaborate with development teams to improve scalability and reliability of services.
- Define and manage SLIs, SLOs, and Error Budgets.
- Lead incident response: organize recovery, document RCA, and run blameless post-mortems.
- Configure and administer Grafana and Zabbix, design insightful dashboards, and fine-tune alerting.
- Integrate and monitor external vendor systems, collaborating with vendor technical support when needed. Key
Requirements:
- Fluent Russian, English B1+ (comfortable with technical documentation).
- 3+ years of experience as an SRE, DevOps, or Infrastructure Engineer.
- Strong understanding of observability principles (metrics, logs, traces).
- Hands-on experience with Grafana and Zabbix (administration, configuration, alert optimization).
- Experience working with AWS and CI/CD tools.
- Practical knowledge of SLI/SLO/Error Budget frameworks.
- Experience leading and documenting incidents and post-mortems.
- Scripting skills for automation (Python, Bash, or Go).
- Solid understanding of distributed systems and networking fundamentals. Nice to Have:
- Experience monitoring and supporting mobile applications.
- Familiarity with Terraform, Prometheus, Loki, ELK, or similar tools.
- Experience working with Kubernetes and containerized environments.
What We Offer
- Fully remote work format.
- Official employment under the Russian Labor Code (for residents of Russia); contractor collaboration available for candidates from other countries.
- Opportunity to work in an international team on a new digital product for the Mexican market.
- A data-driven environment where your contributions have a real impact.
* *[NOVACARD](https://novacard.mx/)**, we’re redefining how people use credit. We are the
* *first interest-free and no-annual-fee credit card in Mexico, designed to simplify personal finances and give users complete control
- all from a mobile app. With NOVACARD, users can access up to
* *$200,000 MXN in credit**, only pay when they use it, and manage everything digitally in under 5 minutes. Our mission is to empower people to make smarter financial decisions by offering flexibility, transparency, and the freedom they need to reach their goals. Simple finances, big goals.*
About the Role
*
* We’re looking for a Site Reliability Engineer (SRE) to ensure the stability, performance, and reliability of our critical production systems. You’ll work at the intersection of development and operations — building automation tools, improving observability, and preventing incidents before they occur.
Key Responsibilities
- Ensure the stability, performance, and fault tolerance of production systems.
- Develop and maintain infrastructure automation and observability tools.
- Monitor system health, respond to incidents, and perform root cause analysis (RCA).
- Collaborate with development teams to improve scalability and reliability of services.
- Define and manage SLIs, SLOs, and Error Budgets.
- Lead incident response: organize recovery, document RCA, and run blameless post-mortems.
- Configure and administer Grafana and Zabbix, design insightful dashboards, and fine-tune alerting.
- Integrate and monitor external vendor systems, collaborating with vendor technical support when needed. Key
Requirements:
- Fluent Russian, English B1+ (comfortable with technical documentation).
- 3+ years of experience as an SRE, DevOps, or Infrastructure Engineer.
- Strong understanding of observability principles (metrics, logs, traces).
- Hands-on experience with Grafana and Zabbix (administration, configuration, alert optimization).
- Experience working with AWS and CI/CD tools.
- Practical knowledge of SLI/SLO/Error Budget frameworks.
- Experience leading and documenting incidents and post-mortems.
- Scripting skills for automation (Python, Bash, or Go).
- Solid understanding of distributed systems and networking fundamentals. Nice to Have:
- Experience monitoring and supporting mobile applications.
- Familiarity with Terraform, Prometheus, Loki, ELK, or similar tools.
- Experience working with Kubernetes and containerized environments.
What We Offer
- Fully remote work format.
- Official employment under the Russian Labor Code (for residents of Russia); contractor collaboration available for candidates from other countries.
- Opportunity to work in an international team on a new digital product for the Mexican market.
- A data-driven environment where your contributions have a real impact.