Logo for XTB online trading

Site Reliability Engineer

Role overview

Qualifications

  • Professional experience in SRE, Infrastructure, or DevOps roles managing high-scale, distributed environments.
  • Advanced programming skills in Python, with a strong focus on building scalable automation, internal tooling, and robust scripts.
  • Hands-on expertise in managing production-grade Kubernetes environments and configuration management tools like Ansible.
  • Proficiency in building standardized telemetry ecosystems using tools like Prometheus and Grafana.

Responsibilities

  • Develop a standardized observability ecosystem with structured events and distributed tracing.
  • Act as a strategic partner to product engineering teams regarding service reliability using error budgets.
  • Enhance detection capabilities to identify issues before they impact the customer through AI/ML.
  • Build internal automation and tooling that streamlines SRE workflows and enhances efficiency.

About the company

XTB online trading logo

XTB online trading

Financial Services

XTB is a global fintech company that provides individual investors with instant access to financial markets from around the world through an innovative online investing platform and the XTB mobile app. Founded in Poland in 2004, we currently support over 1 million customers globally in achieving their investment ambitions. At XTB, we are committed to the ongoing development of the online investing platform enabling our customers to trade 6,300+ instruments including stocks, ETFs, CFDs based on currency pairs, commodities, indices, stocks, ETFs, and cryptocurrencies. With the recent launch of Investment Plans, a long-term passive investing product, our clients can now unlock the growing potential of ETFs and diversify their portfolios effectively. In the key markets, we're offering interest rates on uninvested funds enabling investors to put their money to work and benefit even when they aren't actively investing. Our online platform is a top destination not only for investing but also market analysis and education. We offer an extensive library of educational materials, videos, webinars, and courses to help our customers become better investors irrespective of their trading experience. Our customer service team provides support in 18 languages and is available 24/5 via email, chat, or phone. In over two decades of activity in the financial markets, we have expanded our reach to having more than 1,000 employees. XTB is headquartered in Poland with offices in multiple countries across the globe, including the UK, Germany, Romania, Spain, Czech Republic, Slovakia, Portugal, France, Dubai and Chile. Since 2016, XTB shares have been listed on the Warsaw Stock Exchange. We are regulated by the world’s largest supervisory authorities: Financial Conduct Authority, Polish Financial Supervision Authority, Cyprus Securities & Exchange Commission, Dubai Financial Services Authority and Financial Services Commission. Investing is risky. Invest responsibly.

Company details

IndustryFinancial Services
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

XTB is a global company from the financial industry, focusing on online trading of financial instruments. We are the largest FinTech in Poland and a leader in Central and Eastern Europe, and the range of our operations covers several countries, including Asia and South America. At XTB, we focus on the development of our employees, giving them opportunities to gain knowledge and skills in various fields, as well as offering a number of training and development programs. If you are looking for challenges and want to gain valuable experience in an international business environment, XTB is the right place for you.

We are a certified Great Place to Work company.

We are looking for a Site Reliability Engineer to define and drive the reliability of XTB systems at the scale of millions of clients. In this role, you will strengthen SRE practices and shape the resilience of our entire technology stack through high-impact observability, ensuring our systems remain robust and scalable.


Responsibilities
  • Observability Platform Engineering: Develop a standardized observability ecosystem. Implement a conscious telemetry model focusing on structured events, distributed tracing, and intelligent sampling strategies - that provides deep, actionable insights into system behavior.
  • Reliability Enablement: Act as a strategic partner to product engineering teams, providing the platform, standards, and data they need to own service reliability. Use error budgets and alerting as the primary language for balancing feature velocity with stability.
  • Proactive Resilience & Protection: Enhance detection capabilities to identify issues before they impact the customer. Leverage early-warning systems and AI/ML for automated anomaly detection and intelligent data analysis to continuously verify and strengthen system resilience.
  • Operations & Tooling: Build internal automation and tooling that streamlines SRE workflows, automates routine operational tasks, and enhances efficiency across the technology stack.
  • Incident Management & On-Call Rotation: Participate in an on-call rotation to provide incident management, ensuring rapid incident resolution, effective communication, and post-incident analysis to drive continuous improvement.

  • Requirements
  • Professional Background: Professional experience in SRE, Infrastructure, or DevOps roles managing high-scale, distributed environments.
  • Technical Engineering: Advanced programming skills in Python, with a strong focus on building scalable automation, internal tooling, and robust scripts.
  • Cloud & Orchestration: Hands-on expertise in managing production-grade Kubernetes environments, configuration management tools like Ansible, and designing resilient infrastructure architectures within Azure Kubernetes Service and on-prem environments.
  • Observability Engineering: Proficiency in building standardized telemetry ecosystems. You have mastered self-hosted opensource tools for observability data collection, storage and visualization, like Prometheus, Grafana, ELK Stack, Tempo, Thanos, Jaeger and similar.
  • Operational & Soft Skills: Ability to drive incident management, conduct thorough post-incident analysis, and foster a culture of reliability and shared ownership.

  • Nice to have
  • Experience with commercial APM platforms (e.g., Datadog, Splunk, New Relic) and chaos engineering tooling.
  • Experience with cloud cost management and FinOps principles.
  • Experience defining and tracking SRE metrics (SLI/SLOs) and managing error budgets to drive reliability.
  • Experience with AI/ML techniques for SRE tasks, such as AIOps, automated anomaly detection, log analysis, and optimizing reliability workflows.
  • Experience in building and managing strategies to proactively manage technical debt and align team output with organizational goals.

  • What we offer
  • Real influence on the development of the company and the product.
  • Work in an experienced team that is happy to share its knowledge.
  • A clear vision of development thanks to regular feedback and clear career paths.
  • Regular team-building meetings.

  • Benefits
  • A training budget for courses and conferences that interest you.
  • An extra day off on your birthday.
  • An extra day off for parents.
  • Equipment tailored to your needs.
  • Private medical care and group insurance.
  • Access to an e-learning platform for learning English and a benefits platform.
  • Access to a wellbeing platform and the opportunity to take advantage of workshops and private therapy sessions.
  • Remote work, from the office in Warsaw or from a coworking space in your city.
  • Apply once. Then go straight to the hiring manager.

    After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

    MR

    Marcus Rivera

    Chief Revenue Officer

    m.rivera@company.com
    linkedin.com/in/marcusrivera
    Unlocked after you apply
    ·

    Site Reliability Engineer (SRE) Related jobs

    Other jobs at XTB online trading

    Premium

    Reach out to the hiring manager directly.

    Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

    • Full match report with fit score and gaps
    • Career diagnostics on how recruiters read you
    • Curated company matches and warm intros
    • 48h early access to new roles

    Cancel anytime.