Senior AI-ML Data Scientist
- Strategic Systems International
- Mexico City, Buenos Aires
- MXN 1,200,000 – MXN 1,800,000
Job Summary
We are seeking a Senior AI/ML Data Scientist to own model and agent behavior end to end - from problem framing and algorithm selection through fine-tuning, retrieval design, agentic orchestration, evaluation, and production serving.
This is a hands-on role for someone with genuine depth in machine learning and statistics who is equally comfortable designing an experiment, reading an attention implementation, and shipping the result behind a latency budget. We are particularly interested in candidates who think carefully about agent memory - what an agent should retain, in what form, and how retention is grounded in a governed data warehouse rather than an undifferentiated vector blob.
This role partners closely with the AI Data Engineer, who owns the warehouse, pipelines, and index infrastructure. The boundary: they own the pipeline, the schema, and the guarantees; you own the algorithm, the prompt, and the evaluation.
Key Responsibilities
Machine Learning & Statistical Analysis
Frame ambiguous business problems as tractable ML problems, including the harder prior question of whether ML is the right instrument at all.
Select, implement, and tune models across classical ML and deep learning; know when gradient boosting beats a transformer and say so.
Apply statistical rigor to analysis and experimentation: appropriate hypothesis tests, power analysis, confidence intervals, causal inference where correlation will not carry the claim, and honest treatment of multiple comparisons.
Design and analyze offline evaluations and online experiments (A/B, interleaving, bandits).
Diagnose distribution shift, leakage, and label quality problems before they reach production.
LLM Development & Fine-Tuning
Train, fine-tune, distill, and adapt LLMs using LoRA, QLoRA, DeepSpeed, FSDP, Ray Train, Axolotl, and Hugging Face TRL.
Apply preference optimization and alignment methods (RLHF, PPO, DPO, and successors) where supervised fine-tuning is insufficient.
Generate and validate synthetic data, with explicit attention to the failure modes of training on model output.
Work with multimodal models (vision, text, speech) including LVMs such as LLaVA, Florence-2, and Kosmos-2.
Evaluate frontier and open-weight models - OpenAI, Claude, Gemini, Llama, Mistral, Qwen - and make sourced recommendations on model selection.
Agentic Systems & Memory Architecture
Design and build agents using LangChain, LangGraph, LlamaIndex, CrewAI, the OpenAI Assistants API, or custom orchestrators, with explicit control flow rather than emergent behavior.
Architect agent memory using established cognitive-architecture frameworks - notably CoALA (Cognitive Architectures for Language Agents; Sumers, Yao, Narasimhan & Griffiths). CoALA organizes agents along three axes: memory modules (working, episodic, semantic, procedural), an action space divided into internal actions (retrieval, reasoning, learning) and external grounding actions, and a structured decision-making loop of planning and execution. We expect candidates to be able to critique this framework as well as apply it.
Make and defend concrete memory design decisions: what belongs in episodic versus semantic memory, when memory is written versus consolidated versus forgotten, how procedural memory is updated safely, and how working-memory context budget is allocated under a token limit.
Integrate agentic pipelines with the data warehouse so that agent semantic memory resolves against governed, conformed dimensional data rather than free-floating embeddings - including text-to-SQL over a semantic layer, metric consistency, and enforcement of row-level access controls through the agent boundary.
Design tool interfaces, guardrails, fallback behavior, and human-in-the-loop escalation paths.
Knowledge Graphs & Retrieval
Design ontologies and knowledge graphs; perform entity resolution, relation extraction, and graph construction over unstructured and warehouse sources.
Query and reason over graphs (Neo4j, Neptune, RDF/SPARQL, or property-graph equivalents).
Own retrieval strategy for RAG systems: chunking policy, embedding model selection, hybrid lexical/dense search, reranking, query rewriting, and graph-augmented retrieval measured, not assumed.
Build document intelligence and multimodal retrieval solutions.
Evaluation, Safety & Serving
Build evaluation harnesses that catch regressions before users do: golden datasets, LLM-as-judge with validated agreement to human labels, task-level and end-to-end agent metrics.
Implement guardrails, red-teaming, and safety evaluation; monitor for prompt injection, jailbreaks, and data exfiltration through tool use.
Optimize inference: quantization (GPTQ/INT4/AWQ), pruning, distillation, tensor parallelism, sharding, KV-cache strategy, and token streaming.
Deploy models with vLLM, NVIDIA Triton, Ray Serve, KServe, or TorchServe behind FastAPI or gRPC services, and own the latency, throughput, and cost characteristics of what you ship.
Collaboration
Partner with engineering, product, and client teams to translate business problems into scalable AI solutions.
Communicate uncertainty honestly to non-technical stakeholders.
Mentor junior and mid-level AI/ML engineers.
Required Qualifications
5–10+ years in ML/AI engineering, data science, or related technical roles, with proven experience deploying models at scale in production (LLM, CV, NLP, or multimodal).
ML depth: substantive command of machine learning algorithms and neural network theory -optimization, regularization, attention mechanisms, tokenization, embeddings, and model internals.
Statistics: rigorous grounding in inference, experimental design, and data analysis.
Frameworks: PyTorch (primary), plus TensorFlow or JAX; the Hugging Face ecosystem (Transformers, Datasets, TRL).
Python: expert-level, production-grade. Strong SQL for analysis against a dimensional warehouse.
Agentic systems: production experience with LangChain/LangGraph or equivalent, and a well considered position on agent memory architecture.
Knowledge graphs: hands-on ontology design and graph-based reasoning.
Cloud: expert-level deployment of AI workloads on AWS, Azure, or GCP, including GPU provisioning, cost optimization, containerization, and CI/CD.
Experience with experiment tracking and model lifecycle tooling (MLflow, Weights & Biases).
Preferred Qualifications
Direct experience implementing CoALA or a comparable cognitive architecture (SOAR, ACT R, or a documented in-house framework) in a shipped agent system.
GPU acceleration internals: CUDA, TensorRT, cuBLAS.
Production experience with vLLM, NVIDIA Triton, Ray Serve/Ray Train, DeepSpeed, or FSDP.
Experience with AI security, governance, and compliance frameworks.
Track record of contributing to open-source AI frameworks, or published research.
Ability to lead technical discovery phases and client-facing AI workshops.
Familiarity with lakehouse table formats (Iceberg, Delta Lake) sufficient to collaborate credibly with data engineering.
Skills
- Machine Learning
- Deep Learning
- Statistical Analysis
- Python
- LLM Fine-tuning
- Retrieval-Augmented Generation
- Agentic Orchestration




