Senior SysOps/DevOps Engineer
Détails de l'annonce
Référence
247012
Date de publication
16 septembre 2026
Type de contrat
CDILieu de travail
El Menzah, Tunis, Tunisie
Expérience requise
Entre 2 et 5 ans
Niveau d'études
Bac + 5
Salaire proposé
Disponibilité
Plein temps
Langues
Anglais
Offres similaires
Développeur Applicatif Senior
TUNISIE DIGITAL • Ben Arous
Confirmed QA Engineer
MS SOLUTIONS • Tunis
Administrateur Systèmes & Réseaux
ATEMPO TUNISIE • La Manouba
Responsable Infrastructure et Exploitation IT
ACTIVE CONTACT • Tunis
Ready pour ton next job !!
CPL JOBS • Ariana
Entreprise
CYBERIOUS
Secteur: consulting / étude / stratégie
Taille: Moins de 20 employés
Description de l'annonce
Job Description | Senior MLOps / DevOps Engineer |
Location | This position is remote |
Job Type | Full Time |
Relocation Provided | No |
Compensation | Commensurate with experience |
The Role
We're seeking a Senior SysOps/DevOps Engineer to own platform operations for our AWS infrastructure. You'll be the technical owner of the operations domain: the compute, networking, Kubernetes, and serverless foundations under an AI platform that generates and maintains research on 10,000+ companies, serves institutional clients (including top-tier asset managers), and runs a growing agentic AI workload.
This is a senior, hands-on ownership role on a three-person SysOps/DevOps team — you'll architect and operate the platform alongside teammates focused on CI/CD & security automation and on IAM & internal networking. You'll work directly with the CTO and the AI/full-stack engineering teams, and your work is a direct input to our SOC 2 Type II program and enterprise-client security reviews.
What You'll Work On
Scalable Compute for AI Workloads
Architect and operate auto-scaling execution infrastructure for batch AI generation pipelines (queue-based workers, scale-up/scale-down policies, cost-aware scheduling)
Own our Kubernetes footprint (EKS), including the vector-database layer (Qdrant) powering retrieval at large scale: capacity planning, upgrades, encryption, performance tuning
Own our AWS Lambda architecture for pipeline glue and internal APIs, with cost per invocation tracked and optimized
Support the rollout of managed agent runtimes (Amazon Bedrock AgentCore) — networking, isolation, and observability for agentic AI services
Networking & Security Posture
Own external and internal network architecture: VPC design, NACLs, security groups, private connectivity, VPN/Tailscale access paths
Drive infrastructure hardening for SOC 2 Type II and enterprise security reviews: encryption at rest/in transit, private subnets for data stores, documented access paths
Contribute infrastructure evidence to our compliance tooling and to client/investor technical diligence (architecture docs, security posture)
Reliability, DR & Operations
Own capacity planning and load validation for platform-scale events (library generation at 10k entities plus concurrent client read load), publishing capacity numbers
Design and exercise disaster-recovery: backup strategy, DR tabletops, documented RPO/RTO
Own observability for the infrastructure layer (CloudWatch, dashboards, alerting) and act as senior escalation for infrastructure incidents
Run OS patching as a standing practice: quarterly baseline cadence with Critical ≤7-day / High ≤14-day remediation SLAs, compliance reported monthly
Cost Engineering
Treat infrastructure cost as a first-class metric: instance selection (including Graviton/CPU-optimized fleets for ML inference), storage tiering, right-sizing, and scale-down discipline
Partner with the AI team on the unit economics of generation and retrieval — infrastructure choices that cut cost-per-entity without sacrificing reliability
Requirements
Essential
4+ years of SysOps/DevOps/SRE experience, with 2+ years operating production AWS at platform-ownership level
Deep AWS expertise: EC2, EKS, Lambda, RDS, S3, VPC networking, IAM fundamentals, CloudWatch
Production Kubernetes experience: cluster operations, upgrades, capacity planning, stateful workloads
Strong networking fundamentals: VPC design, security groups/NACLs, VPN/private connectivity, TLS
Proven auto-scaling and queue-based architecture experience for batch or high-throughput workloads
Security-first operational mindset: patching discipline, encryption, least-privilege, audit-ready documentation
Scripting/automation proficiency (Python and/or Bash)
Fluent English communication skills (written and verbal)
Highly Valued
Experience supporting SOC 2 (Type I/II), ISO 27001, or similar compliance programs — evidence collection, control implementation, auditor interaction
Vector database operations (Qdrant, Weaviate, Milvus) or other stateful data infrastructure on Kubernetes
Experience running infrastructure for ML/AI workloads (GPU/CPU inference fleets, batch pipelines, model serving, Bedrock/SageMaker)
Distributed PostgreSQL operations (replication, pgEdge/Aurora, migration off public subnets, zero-downtime changes)
Disaster-recovery design and testing (RPO/RTO definition, tabletop exercises)
Cost-optimization track record with measurable results (FinOps practices, Graviton adoption, savings plans)
Exposure to compliance tooling and security scanning pipelines
Experience in fintech, or other environments with enterprise security review processes
Technical Stack You'll Use
Cloud: AWS (EC2, EKS, Lambda, RDS/PostgreSQL, S3, CloudWatch, SQS, Bedrock)
Containers & Orchestration: Kubernetes (EKS), Docker
Data Infrastructure: Qdrant (vector DB on EKS), PostgreSQL/pgEdge, Neo4j
Networking & Access: VPC, Tailscale/VPN, Auth0 (SSO/SAML integration points)
Compliance & Security: compliance tooling, dependency/secret scanning pipelines, CVE remediation workflows
Automation: Python, Bash, CI/CD (GitHub Actions / AWS CodeBuild)
AI Platform (what you'll support): LangGraph agentic workflows, Bedrock AgentCore, ONNX/CPU inference fleets
What We Offer
High Impact: Own the operations domain outright — your architecture decisions carry the platform's two hardest goals: scale and unit cost
Autonomy: Ownership of infrastructure from design to production operation, working directly with the CTO
Modern Stack: Infrastructure for cutting-edge agentic AI at enterprise-grade security standards
Remote First: Work from anywhere with strong English communication
Learning: Rapid exposure to AI-platform operations, compliance engineering, and enterprise fintech infrastructure
Personal Attributes
Self-Directed: Thrive with minimal supervision, define your own milestones and deliver
Pragmatic: Balance perfection with shipping; incremental infrastructure improvement over big-bang re-architecture
Calm Under Fire: Systematic incident response; root causes over quick patches; blameless postmortems
Security-Minded: Treat auditability and least-privilege as defaults, not chores
Collaborative: Work effectively across AI, full-stack, and QA engineering in a distributed, multi-time-zone team
Efficient: Small team means high productivity and smart tradeoffs are essential
To Apply
Please include:
Resume highlighting relevant platform/infrastructure ownership experience
Brief description of your most impactful infrastructure project (what you built or scaled, the constraint you were solving for — cost, scale, or compliance — and the measurable result)
Link to GitHub/portfolio if available (optional but helpful)