XD logo

Staff Platform Engineer

Abu Dhabi, United Arab Emirates
Full time
On-site

Job description

Apply now
Role Overview
We are looking for a Staff Platform Engineer to define and own the backend platform infrastructure that the AI Factory’s engineering teams build on. This is a staff-level individual contributor role: you set technical direction, make architectural decisions that hold across multiple teams and years, and operate as a force multiplier for every engineer who depends on what you build.
This is a backend-first role. Your focus is the platform — not the AI. You will own the services, APIs, event architecture, data pipelines, deployment systems, and operational tooling that keep production running reliably at government scale. You do not need to be an AI specialist. You do need to be an expert platform and backend engineer with the depth and range to architect across distributed systems, event-driven infrastructure, cloud operations, and database reliability, and the judgment to make the right calls without being directed.
At staff level, we expect more than excellent execution. You will shape how the platform is built, set standards that others follow, surface and resolve systemic risks before they become incidents, and ensure the platform evolves ahead of what the teams building on it will need next. You are the technical anchor for platform quality and reliability across the AI Factory.
Core Responsibilities
•       Define the architecture of the AI Factory’s backend platform: service boundaries, API contracts, inter-service communication patterns, and the structural decisions that determine how well the platform holds up as it scales across multiple government entities.
•       Architect, build, and operate the event-driven infrastructure that underpins the AI Factory’s systems: event streaming pipelines, message broker design (Kafka, Azure Service Bus, Azure Event Hubs, or equivalent), consumer group topology, event schema governance, and the reliability patterns that make asynchronous systems trustworthy in production.
•       Own the data infrastructure end-to-end: ingestion pipelines, transformation layers, PostgreSQL at scale, object storage, caching, and the ETL processes that keep data current, consistent, and trustworthy across multiple consumers.
•       Design and operate the deployment platform: containerised service infrastructure, Kubernetes operations, CI/CD pipelines, infrastructure-as-code, and the release engineering practices that let teams ship safely and frequently.
•       Set and own the observability standard across the platform: structured logging, distributed tracing, metrics design, SLO definition and tracking, error budget management, alerting strategy, and the tooling that makes system behaviour transparent.
•       Own cloud infrastructure architecture on Azure: compute, networking, storage, managed services, access control, cost management, and the cross-cutting decisions that apply to every service the AI Factory runs.
•       Lead platform reliability engineering: failure mode analysis, capacity planning, SLO/SLA ownership, chaos engineering practices, and the systematic work that prevents incidents rather than just responding to them.
•       Lead incident response for platform-level issues: triage, resolution, post-mortem facilitation, and the follow-through that turns incidents into durable systemic improvements.
•       Build shared platform abstractions and internal tooling that genuinely raise engineering productivity — and take responsibility for their architecture, adoption, quality, and long-term evolution.
•       Set platform engineering standards across the AI Factory: coding standards, API design conventions, infrastructure-as-code practices, security requirements, and the technical bar that applies to everything that runs in production.
•       Partner to shape the overall platform roadmap and ensure backend infrastructure evolves in step with product and AI engineering needs.
•       Own platform security architecture: secrets management, network segmentation, identity and access management, zero-trust patterns, and compliance with government security and data sovereignty requirements.
 
Basic Qualifications
•       10+ years of experience in backend or platform engineering, with a demonstrable track record of making architectural decisions at scale that have held up over time — not just contributing to systems built by others.
•       Deep expertise in event-driven architecture: designing and operating event streaming systems (Kafka, Azure Service Bus, Azure Event Hubs, or equivalent) at production scale, including schema evolution, consumer topology, exactly-once semantics, and the operational discipline required to run asynchronous infrastructure reliably.
•       Expert-level backend engineering in one or more of Python, Java, or Go — with the depth to make sound architectural decisions, set coding standards, and evaluate trade-offs across distributed systems, not just write good code.
•       Expert-level knowledge of cloud infrastructure on Azure — architecture, compute, networking, storage, managed services, cost optimisation, and security — at the depth required to set the standard for the entire organisation. AWS or GCP experience at equivalent depth also valued.
•       Deep expertise in containerisation and orchestration: Kubernetes architecture, cluster operations, networking, storage, resource management, and the operational depth to run complex containerised workloads reliably at scale.
•       Expert-level knowledge of PostgreSQL at production scale: schema design, query optimisation, indexing strategies, connection pooling, replication, failover, and the operational discipline to run a critical relational database without constant intervention.
•       Deep expertise in distributed systems design: consistency models, distributed transactions, failure modes, idempotency, back-pressure, and the trade-offs between synchronous and asynchronous communication patterns.
•       Expert-level observability: SLO design, error budget management, distributed tracing architecture, metrics strategy, alerting design, and the ability to build a platform that surfaces its own health without requiring manual investigation.
•       Demonstrated ability to set technical direction and standards across a multi-team engineering organisation — and to drive adoption of those standards through influence, documentation, and technical credibility.
•       Strong written and verbal communication — you can author technical RFCs, drive architectural decisions in group settings, and represent complex infrastructure trade-offs clearly to engineering leaders and non-technical stakeholders.
 
Preferred Qualifications
•       Experience designing and operating a platform-as-a-product — with a track record of building internal developer tooling that is genuinely adopted, trusted, and treated as a product in its own right.
•       Experience with service mesh architecture (Istio, Linkerd, or similar) at production scale, including mTLS, traffic management, observability integration, and the operational overhead of running a mesh in a multi-team environment.
•       Experience with CQRS, event sourcing, or saga patterns at production scale — including the trade-offs, failure modes, and operational complexity these patterns introduce in real systems.
•       Experience with advanced data infrastructure: multi-model data stores, polyglot persistence, caching layer design (Redis), search infrastructure (Elasticsearch, OpenSearch), or time-series data at scale.
•       Familiarity with the infrastructure demands of AI workloads at the platform layer — not as an AI engineer, but with enough architectural depth to ensure the backend platform supports retrieval pipelines, embedding infrastructure, and model-serving layers without becoming a bottleneck or single point of failure.
•       Experience in a regulated or government-adjacent environment: data sovereignty requirements, audit trails, compliance frameworks, and the engineering discipline required to operate sensitive data infrastructure to a high standard.
•       Experience with platform security at depth: zero-trust networking, secrets rotation automation, identity federation across multiple systems, and the proactive security practices that keep a high-value infrastructure target hardened over time.
 
Our Stack
We use the tools best suited to each problem. Current defaults reflect what works well for our use cases — not a mandated standard. At staff level, you will influence what this list looks like over time.
•       Backend services:Python, Java, REST APIs, gRPC, WebSockets / SSE
•       Event architecture:Apache Kafka, Azure Service Bus, Azure Event Hubs, schema registry
•       Data:PostgreSQL, Redis, object storage (Azure Blob), document ingestion pipelines, ETL tooling
•       Infrastructure:Azure, Docker, Kubernetes, Terraform, CI/CD pipelines
•       Observability:Structured logging, distributed tracing, Prometheus / Grafana or equivalent, SLO tooling, alerting
•       Security:Azure AD / Entra ID, secrets management, network policy, RBAC, zero-trust, identity federation
 
How We Work
•       Ownership without scope limits. Takes responsibility for the platform from architecture through operation. Does not wait for a ticket. Identifies what needs to exist, builds it, and keeps it running.
•       Force multiplier.Measures success by the reliability and velocity it enables in the teams that depend on the platform — not just the systems it directly builds. The best platform work disappears into the background.
•       Evidence-driven.Grounds every architectural decision in evidence: load profiles, failure data, cost analysis, and real operational experience. Knows the difference between a good abstraction and a premature one, and can defend the distinction.
•       Reliability by design.Thinks in failure modes from the first line of architecture. Designs systems to fail safely, proposes SLOs rather than waiting to be asked, and drives the systemic work that prevents incidents before they happen.
•       Raises the bar.Sets a technical standard that others orient to. Authors RFCs, runs design reviews, mentors, and documents in a way that transfers depth — not just solves immediate problems.
•       Clear thinker.Writes and communicates with precision. Can explain complex architectural decisions and distributed systems trade-offs to engineers across specialisations, and to stakeholders who do not share their technical background.
 
Technical Depth Expectations
Candidates will be expected to demonstrate genuine depth in at least four of the following areas. At staff level, depth means the ability to set architectural direction, define standards, and make decisions that shape how an entire engineering organisation operates — not just execute well independently.
•       Event-driven architecture — event streaming design, message broker operations, schema evolution, consumer topology, exactly-once delivery, saga and outbox patterns, and the failure modes of asynchronous distributed systems at scale.
•       Distributed systems design — consistency models, CAP trade-offs, distributed transactions, idempotency, back-pressure, circuit breaking, and the architectural decisions that determine how a platform behaves when things go wrong.
•       Backend service architecture — API design, service decomposition, inter-service communication patterns, and the structural decisions that keep a multi-service platform coherent, evolvable, and scalable.
•       Cloud infrastructure at scale — architecture, networking, compute, storage, and cost management at the depth required to set organisational standards and make decisions that hold over time on Azure or equivalent.
•       Data infrastructure design — pipeline architecture, polyglot persistence, schema design, reliability patterns, and the engineering discipline to keep complex data infrastructure trustworthy across multiple consumers.
•       Observability and reliability engineering — SLO design, error budget management, instrumentation architecture, distributed tracing, alerting strategy, and the practices that turn operational data into systemic improvements.
•       Platform security architecture — zero-trust design, identity federation, secrets management, network segmentation, and the proactive security practices appropriate for high-value government infrastructure.
Apply for this job
View all jobs