AI + Reliability + Cloud + Backend

Zunaira Siddique

Senior engineer focused on AI systems, site reliability, cloud platforms, and software delivery. I build resilient infrastructure, practical automation, and backend systems that stay dependable as products grow.

What I Optimize For Reliable systems, faster delivery, calmer operations.
AI Platforms

Applied AI infrastructure, integrations, and production-minded delivery.

SRE

Observability, incident readiness, operational resilience, and service trust.

Cloud

Multi-cloud architecture, platform engineering, automation, and security-aware operations.

Software

Go and Python services, APIs, tooling, and delivery workflows that teams can scale with.

About

Engineering that balances product speed with operational depth.

I bring more than ten years of hands-on experience across DevOps, cloud-native engineering, backend systems, platform reliability, LLMOps, and modern automation. My work spans Kubernetes, Docker, AWS, GCP, observability tooling, CI/CD, and software development in Go and Python.

I care about systems that work well in the real world: clear delivery paths, strong operational visibility, thoughtful developer experience, and infrastructure that supports teams as products and complexity grow.

What I Bring

Four areas that define how I help teams move with confidence.

Production reliability

Stronger observability, cleaner incident response, and systems designed to fail more gracefully under real operating pressure.

Platform clarity

Better internal platforms, streamlined workflows, and infrastructure patterns that reduce friction for engineering teams.

AI-ready foundations

Practical infrastructure and software patterns that make AI initiatives easier to integrate, operate, and scale responsibly.

Delivery confidence

Automation, release discipline, and backend engineering that keep shipping predictable without sacrificing quality.

Some problems I was brought in to solve, and what shipped.

Okta onboarding automation

Problem

Onboarding a new application to Okta was an 8-step manual process that slowed platform adoption.

Action

Automated the workflow with Python, AI agents, GitHub Actions, and automation to guide and execute onboarding steps.

Result

Cut the process from 8 steps to 3, removing a recurring manual bottleneck for onboarding teams and helping onboard 200+ applications on a tight deadline.

Zero-downtime Open edX platform

Problem

Open edX needed scalable, secure, and resilient deployments across development, staging, and production on AWS.

Action

Led the team automating deployments with Ansible, Terraform, and CloudFormation, stood up the Tutor distribution on EKS, and configured Jenkins on ECS and Azure DevOps.

Result

Delivered zero-downtime, easy-rollback continuous deployment across every environment.

Multi-cloud Kubernetes automation

Problem

Deploying Kubernetes and Istio components consistently across multiple cloud providers was manual and error-prone.

Action

Built gRPC services in Golang to automate deployment of Kubernetes and Istio components across AWS, GCP, and DigitalOcean.

Result

Standardized multi-cloud infrastructure deployment behind a single automated interface.

AWS to GCP migration

Problem

A production React application needed to move from AWS to GCP without losing monitoring coverage.

Action

Executed the migration using GCP's migration service and stood up an EFK stack for monitoring and logging on the new platform.

Result

Completed the cloud migration with logging and monitoring in place from day one.

ECS/RDS to EKS/Aurora migration

Problem

An application running on ECS with RDS needed to move onto Kubernetes with GitOps-driven delivery, without disrupting production traffic.

Action

Migrated the workload from ECS and RDS to EKS and Aurora, then configured ArgoCD to manage deployments declaratively from Git.

Result

Moved the workload onto EKS and Aurora with zero production disruption, and replaced manual deploys with Git-driven releases through ArgoCD.

GitOps rollout automation for NodeJS microservices

Problem

NodeJS-hosted microservices needed automated workflows and progressive rollouts, but nothing tied deployment orchestration directly into the Go tooling already in use.

Action

Implemented Argo Workflows, ArgoCD, and Argo Rollouts using the Golang SDK, and wrote a custom Kubernetes CRD to model the rollout process natively.

Result

Gave the microservices GitOps-driven, progressive delivery with rollout logic controlled through Kubernetes itself rather than external scripts.

Skills

Tools and technologies I use across platform and product work.

AWS Azure GCP DigitalOcean Amazon Bedrock AgentCore Guardrails SageMaker AI Integrations LLM Application Infrastructure RAG Vector DBs LangGraph Golang Python FastAPI Docker Kubernetes Helm Okta CI/CD Jenkins GitLab CI GitHub Actions Monitoring and Observability Elasticsearch Prometheus Grafana Kibana Terraform Ansible Linux SQL MySQL REST APIs gRPC HashiCorp Vault Playwright

Certifications

Verified credentials, rendered locally for reliability.

Experience

A path shaped by platform, product, and reliability work.

Senior Software Engineer

Zalando

Oct 2023 - Oct 2024

Senior DevOps Engineer

Arbisoft

Dec 2020 - Oct 2023

Senior Software Engineer

Acquia

Jul 2022 - Dec 2022

Software Engineer

Cloudplex

Jul 2019 - Dec 2020

DevOps Engineer

Time Xcess

Jan 2016 - Jun 2019

Contact

Let’s build something reliable, efficient, and well-crafted.

Reach me directly at zunairasiddique518@gmail.com for collaborations, consulting, and engineering opportunities.

The fastest path is a direct email or a calendar booking if you already know what you want to build or improve.