Other

Sr. AI Operations Engineer

RemoteπŸ‡¦πŸ‡· πŸ‡§πŸ‡· πŸ‡¨πŸ‡΄ πŸ‡¨πŸ‡· πŸ‡²πŸ‡½ πŸ‡ΊπŸ‡ΎArgentina, Brazil, Colombia, Costa Rica, Mexico, Uruguay

The engineer will work with Databricks and Microsoft Copilot ecosystems, support model routing and gateways, build observability and FinOps controls, and implement agent security and lifecycle processes. This is an opportunity to shape repeatable operational patterns that enable safe, cost-effective, and scalable AI for thousands of enterprise users.

Stack

Generative AIAI Agents & Agent FrameworksLLMsRAG ArchitecturesPythonREST APIsAPI SecuritySecrets ManagementCloud IdentityAI PlatformsKubernetesTerraform

Summary

We are seeking a Senior AI Operations (AIOps) Engineer to operate, secure, monitor, and optimize an enterprise Generative AI and agent platform. This role focuses on production operation and governance of AI agents, model gateways, and supporting services to ensure reliability, security, efficiency, and compliance with enterprise standards. The position sits at the intersection of AI engineering, platform engineering, security, observability, and FinOps and emphasizes operational excellence over model development.

The engineer will work with Databricks and Microsoft Copilot ecosystems, support model routing and gateways, build observability and FinOps controls, and implement agent security and lifecycle processes. This is an opportunity to shape repeatable operational patterns that enable safe, cost-effective, and scalable AI for thousands of enterprise users.

Responsibilities

  • Operate and support enterprise AI agents across Databricks and Microsoft Copilot platforms, including deployment, monitoring, versioning, and retirement.
  • Establish and enforce standards for agent configuration, lifecycle management, and promotion from experimentation to governed production.
  • Build and maintain observability for AI applications, model endpoints, agents, and AI gateways, including dashboards and alerts.
  • Monitor model and agent latency, token consumption, request volume, error rates, tool/function-call failures, retrieval/RAG performance, and agent execution traces.
  • Lead incident response and root-cause analysis for AI platform and agent incidents; implement remediation and preventive actions.
  • Design and optimize context strategies: system prompts, context windows, retrieval patterns, memory strategies, and tool selection to balance quality, latency, and cost.
  • Monitor and optimize token consumption and develop mechanisms for usage attribution, showback/chargeback, budgeting, and alerting.
  • Implement and enforce AI and agent security controls, apply least-privilege principles, and mitigate risks such as prompt injection, data leakage, unauthorized tool execution, and credential exposure.
  • Support centralized AI gateway and model operations: configure model routing, fallbacks, rate limits, quotas, authentication, and access controls for internal and external models.
  • Collaborate with platform, security, finance, and MLOps teams to forecast consumption, integrate governance controls, and operationalize Databricks AI and Microsoft Copilot capabilities.
  • Support the evolution toward multi-agent and agent-to-agent architectures as enterprise adoption matures.

Requirements

  • 4+ years of experience in cloud, platform, DevOps, SRE, MLOps, or AI engineering.
  • Hands-on experience supporting Generative AI or LLM-based applications and AI agents or agent frameworks.
  • Strong understanding of LLM inference, tokens and context windows, prompt and context engineering, RAG architectures, embeddings and vector search, tool/function calling, and model routing/fallback.
  • Experience building enterprise observability and monitoring for model endpoints, agents, and gateways, including dashboards and alerts.
  • Experience with cloud identity, RBAC, secrets management, service principals, and API security (OAuth 2.0, OIDC, managed identities, and workload identities as applied to AI agents and workloads).
  • Strong Python and scripting skills and experience working with REST APIs and API-based platforms.
  • Experience with CI/CD and infrastructure automation and strong troubleshooting skills across distributed systems, with a strong DevOps/platform engineering foundation across production deployment, reliability, monitoring, and cloud infrastructure.
  • Hands-on experience operating enterprise AI platforms, preferably within Databricks and/or Microsoft Copilot ecosystems.
  • Experience with Kubernetes-based production environments, containerized workloads, and cloud-native deployment patterns; AKS experience is preferred.
  • Experience operating AI/model gateways, including model routing, fallback, rate limiting, quotas, authentication, and observability.

Nice to Have

  • Experience with Terraform or similar Infrastructure-as-Code (IaC) tools
  • Experience with LiteLLM, vLLM, or similar AI gateway/model-serving technologies
  • Experience with observability frameworks like OpenTelemetry and vector databases for enterprise RAG architectures
  • Experience with multi-agent or agent-to-agent architectures
Other

Sr. AI Operations Engineer

Location
Remote
Hiring in
Argentina, Brazil, Colombia, Costa Rica, Mexico, Uruguay
Compensation
USD
I'm interested!← Back to all roles

Apply now

Interested in this role?

Send your resume and a brief introduction to jobs@techwarely.com and we'll get back to you soon.

Send us an email