AI/ML Engineer – Agentic Systems

This Position Is Closed
Python, PyTorch, Hugging Face ecosystem, MLOps, LLM applications
Location:
Remote
: 
, 
Argentina 🇦🇷 / Brazil 🇧🇷 / Colombia 🇨🇴

About the role

We're hiring an ML engineer with deep modeling fundamentals who now builds LLM-powered agents and the systems that prove they work. You'll own agentic workflows end to end: design, evaluation, tracing, and production monitoring.

This is not a prompt-tinkering role. We want someone who brings experimental rigor from classical and deep ML (baselines, ablations, held-out sets, error analysis) to non-deterministic, multi-step LLM systems.

Who are we looking for?

Responsibilities

  • Design and ship LLM-based agents and multi-step workflows: tool use, planning, retrieval (RAG), memory, and multi-agent orchestration.
  • Make deliberate architecture choices and integrate agents with internal data and APIs via function calling and MCP, with guardrails, fallbacks, and human-in-the-loop checkpoints.
  • Take systems from prototype to production, owning latency, cost, and reliability budgets.
  • Define what "good" means for each use case with business stakeholders, define eval criteria and build eval suites
  • Develop and calibrate LLM-as-judge graders against human labels; measure judge agreement and bias.
  • Evaluate both final outputs and trajectories and run evals in CI so prompt, model, and tool changes are gated on measured quality, not vibes.
  • Instrument agents with structured tracing (spans for LLM calls, tool calls, retrieval).
  • Monitor quality, hallucination and grounding rates, latency, token cost, and failure modes on live traffic.
  • Run systematic error analysis on traces, detect drift and regressions when upstream models or data change, fine-tune or distill models (LoRA, SFT, preference tuning).
  • Train classical or deep models where they beat an LLM on cost, latency, or accuracy 
  • Design experiments with proper statistics, set team standards for evals and observability
  • Explain system behavior, eval results, and trade-offs clearly to non-technical leaders.

Skills & Experience

  • Degree in Computer Science or a related field; Master's or PhD in ML is a plus.
  • 5+ years of hands-on ML engineering, with models you trained, evaluated, and deployed to production
  • Strong grasp of the fundamentals: generalization, overfitting and leakage, evaluation metrics, calibration, and experimental design.
  • Experience with PyTorch and the Hugging Face ecosystem; working understanding of transformer internals, embeddings, and tokenization.
  • Experience with MLOps
  • 2+ years building LLM applications, including agentic or multi-step tool-using systems.
  • Hands-on with at least one agent framework (LangGraph, Pydantic AI, OpenAI Agents SDK, or similar) and a clear view of when not to use one.
  • Working knowledge of modern RAG design.
  • Experience with structured outputs, function calling, and prompt and context engineering for frontier models (Claude, GPT, Gemini) and open-weight models.
  • Experience in building eval framework or suite from scratch, not just run benchmarks.
  • Experience with LLM tracing and eval tooling such as Langfuse, LangSmith, Arize Phoenix, Braintrust, or Weights & Biases Weave.
  • Advanced Python and production-quality code: typing, testing, async, and API design.
  • Strong SQL for analyzing traces, logs, and large datasets.
  • Cloud deployment experience, containers, and CI/CD.

Nice to have

  • Background in AI safety, red-teaming, or guardrails (prompt injection, PII leakage, jailbreak testing).
  • Experience in financial services or other regulated domains where auditability and traceability matter.
  • Publications, open-source contributions, or public write-ups on evals, agents, or ML systems.
  • Comfort working in ambiguous, fast-moving environments and taking ideas from PoC to production with light guidance.

What we offer

Work:

  • Flexible working hours;
  • Collaborative, friendly team environment;
  • Remote/Hybrid work;

Life:

  • Company social events;
  • Annual corporate parties;

Health:

  • Comprehensive medical insurance;

Education:

  • Allowances for professional education;
  • English language courses with native speakers;
  • Internal knowledge-sharing sessions.

What we offer

About Proxet

Proxet is a professional software development firm trusted by clients from around the world. With our expertise in AI and machine learning, we help businesses reimagine their possibilities and transform ideas into tangible digital solutions. By providing core services with an emphasis on data practices, we shape the future, one step at a time.

If you’d like to join our Proxet Nation and work closely with high-level professionals and our engineers, fill in the form!

RECRUITER:

Anastasiia Nesterenko

Interested? Let's get in touch!

Tell us about yourself, then leave a link or upload your resume and we will get back to you soon!

Max file size 10MB.
Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.

💌

Thank you!

We received your application and our recruiters will contact you as soon as possible!
Oops! Something went wrong while submitting the form.

Interested in this closed position? Let's get in touch!

Tell us about yourself, then leave a link or upload your resume and we will get back to you soon!

Max file size 10MB.
Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.

💌

Thank you!

We received your application and our recruiters will contact you as soon as possible!
Oops! Something went wrong while submitting the form.

Are you already working at Proxet and do you have a friend who's a good fit for us?

Tell us about your friend - send their resume to our email careers@proxet.com!