Senior SysOps/DevOps Engineer

CYBERIOUS El Menzah, Tunis Publiée le 16 septembre 2026

Détails de l'annonce

Référence

247012

Date de publication

16 septembre 2026

Type de contrat

CDI

Lieu de travail

El Menzah, Tunis, Tunisie

Expérience requise

Entre 2 et 5 ans

Niveau d'études

Bac + 5

Salaire proposé

2000 - 3000 TND / Mois

Disponibilité

Plein temps

Langues

Anglais

Offres similaires

Développeur Applicatif Senior

TUNISIE DIGITAL Ben Arous

Confirmed QA Engineer

MS SOLUTIONS Tunis

Administrateur Systèmes & Réseaux

ATEMPO TUNISIE La Manouba

CPL JOBS  logo

Ready pour ton next job !!

CPL JOBS Ariana

Entreprise

CYBERIOUS

Secteur: consulting / étude / stratégie

Taille: Moins de 20 employés

Description de l'annonce


Job Description

Senior MLOps / DevOps Engineer

Location

This position is remote

Job Type

Full Time

Relocation Provided

No

Compensation

Commensurate with experience

The Role

We're seeking a Senior SysOps/DevOps Engineer to own platform operations for our AWS infrastructure. You'll be the technical owner of the operations domain: the compute, networking, Kubernetes, and serverless foundations under an AI platform that generates and maintains research on 10,000+ companies, serves institutional clients (including top-tier asset managers), and runs a growing agentic AI workload.

This is a senior, hands-on ownership role on a three-person SysOps/DevOps team — you'll architect and operate the platform alongside teammates focused on CI/CD & security automation and on IAM & internal networking. You'll work directly with the CTO and the AI/full-stack engineering teams, and your work is a direct input to our SOC 2 Type II program and enterprise-client security reviews.

What You'll Work On

Scalable Compute for AI Workloads

  • Architect and operate auto-scaling execution infrastructure for batch AI generation pipelines (queue-based workers, scale-up/scale-down policies, cost-aware scheduling)

  • Own our Kubernetes footprint (EKS), including the vector-database layer (Qdrant) powering retrieval at large scale: capacity planning, upgrades, encryption, performance tuning

  • Own our AWS Lambda architecture for pipeline glue and internal APIs, with cost per invocation tracked and optimized

  • Support the rollout of managed agent runtimes (Amazon Bedrock AgentCore) — networking, isolation, and observability for agentic AI services

Networking & Security Posture

  • Own external and internal network architecture: VPC design, NACLs, security groups, private connectivity, VPN/Tailscale access paths

  • Drive infrastructure hardening for SOC 2 Type II and enterprise security reviews: encryption at rest/in transit, private subnets for data stores, documented access paths

  • Contribute infrastructure evidence to our compliance tooling and to client/investor technical diligence (architecture docs, security posture)

Reliability, DR & Operations

  • Own capacity planning and load validation for platform-scale events (library generation at 10k entities plus concurrent client read load), publishing capacity numbers

  • Design and exercise disaster-recovery: backup strategy, DR tabletops, documented RPO/RTO

  • Own observability for the infrastructure layer (CloudWatch, dashboards, alerting) and act as senior escalation for infrastructure incidents

  • Run OS patching as a standing practice: quarterly baseline cadence with Critical ≤7-day / High ≤14-day remediation SLAs, compliance reported monthly

Cost Engineering

  • Treat infrastructure cost as a first-class metric: instance selection (including Graviton/CPU-optimized fleets for ML inference), storage tiering, right-sizing, and scale-down discipline

  • Partner with the AI team on the unit economics of generation and retrieval — infrastructure choices that cut cost-per-entity without sacrificing reliability

Requirements

Essential

  • 4+ years of SysOps/DevOps/SRE experience, with 2+ years operating production AWS at platform-ownership level

  • Deep AWS expertise: EC2, EKS, Lambda, RDS, S3, VPC networking, IAM fundamentals, CloudWatch

  • Production Kubernetes experience: cluster operations, upgrades, capacity planning, stateful workloads

  • Strong networking fundamentals: VPC design, security groups/NACLs, VPN/private connectivity, TLS

  • Proven auto-scaling and queue-based architecture experience for batch or high-throughput workloads

  • Security-first operational mindset: patching discipline, encryption, least-privilege, audit-ready documentation

  • Scripting/automation proficiency (Python and/or Bash)

  • Fluent English communication skills (written and verbal)

Highly Valued

  • Experience supporting SOC 2 (Type I/II), ISO 27001, or similar compliance programs — evidence collection, control implementation, auditor interaction

  • Vector database operations (Qdrant, Weaviate, Milvus) or other stateful data infrastructure on Kubernetes

  • Experience running infrastructure for ML/AI workloads (GPU/CPU inference fleets, batch pipelines, model serving, Bedrock/SageMaker)

  • Distributed PostgreSQL operations (replication, pgEdge/Aurora, migration off public subnets, zero-downtime changes)

  • Disaster-recovery design and testing (RPO/RTO definition, tabletop exercises)

  • Cost-optimization track record with measurable results (FinOps practices, Graviton adoption, savings plans)

  • Exposure to compliance tooling and security scanning pipelines

  • Experience in fintech, or other environments with enterprise security review processes

Technical Stack You'll Use

  • Cloud: AWS (EC2, EKS, Lambda, RDS/PostgreSQL, S3, CloudWatch, SQS, Bedrock)

  • Containers & Orchestration: Kubernetes (EKS), Docker

  • Data Infrastructure: Qdrant (vector DB on EKS), PostgreSQL/pgEdge, Neo4j

  • Networking & Access: VPC, Tailscale/VPN, Auth0 (SSO/SAML integration points)

  • Compliance & Security: compliance tooling, dependency/secret scanning pipelines, CVE remediation workflows

  • Automation: Python, Bash, CI/CD (GitHub Actions / AWS CodeBuild)

  • AI Platform (what you'll support): LangGraph agentic workflows, Bedrock AgentCore, ONNX/CPU inference fleets

What We Offer

  • High Impact: Own the operations domain outright — your architecture decisions carry the platform's two hardest goals: scale and unit cost

  • Autonomy: Ownership of infrastructure from design to production operation, working directly with the CTO

  • Modern Stack: Infrastructure for cutting-edge agentic AI at enterprise-grade security standards

  • Remote First: Work from anywhere with strong English communication

  • Learning: Rapid exposure to AI-platform operations, compliance engineering, and enterprise fintech infrastructure

Personal Attributes

  • Self-Directed: Thrive with minimal supervision, define your own milestones and deliver

  • Pragmatic: Balance perfection with shipping; incremental infrastructure improvement over big-bang re-architecture

  • Calm Under Fire: Systematic incident response; root causes over quick patches; blameless postmortems

  • Security-Minded: Treat auditability and least-privilege as defaults, not chores

  • Collaborative: Work effectively across AI, full-stack, and QA engineering in a distributed, multi-time-zone team

  • Efficient: Small team means high productivity and smart tradeoffs are essential

To Apply

Please include:

  • Resume highlighting relevant platform/infrastructure ownership experience

  • Brief description of your most impactful infrastructure project (what you built or scaled, the constraint you were solving for — cost, scale, or compliance — and the measurable result)

  • Link to GitHub/portfolio if available (optional but helpful)

Partager: