AI Support Engineer
About Millennium
Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millennium’s mission is to deliver results for our investors.
Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact.
Meet the Team
Core to the health and growth of Millennium’s business, the Information Technology organization develops flexible, scalable technology and advanced proprietary systems, including the development of the next generation of analytical and trading capabilities. Within this organization, the Infrastructure team supports critical platforms and services across the firm, with a focus on reliability, operational excellence, and strong end-user trust in rapidly evolving technology environments.
What You'll Do
- Serve as the first line of response for customer and application issues across the firm’s AI platforms
- Instrument, maintain, and improve monitoring and alerting for AI services, including LLM and agent monitoring frameworks
- Implement and strengthen incident management and postmortem processes to improve operational reliability
- Resolve complex technical issues end to end by applying strong troubleshooting skills and AI domain expertise
- Escalate issues to engineering teams when appropriate, with clear context and effective follow-through
- Author and maintain support processes, documentation, and user guides
- Partner with QA and Engineering teams to improve the quality, stability, and adoption of AI services
- Continuously improve support processes, tooling, and methodologies, and participate in a global follow-the-sun 24x7 on-call rotation
What You Bring
- Proven experience supporting AI platforms and working with OpenAI, Anthropic, and other model vendors
- Hands-on experience with retrieval-augmented generation and vector databases
- Deep experience configuring, administering, and troubleshooting Cursor and Claude Code
- Proficiency in Python or Java, with the ability to code, debug, and solve technical issues directly
- Experience supporting and maintaining FastAPI-based services
- Familiarity with Linux and container runtime environments such as Kubernetes
- Experience building or improving DevOps release pipelines, alerting processes, and monitoring using tools such as Rootly, Datadog, LangSmith, or Pydantic
- Strong analytical, problem-solving, communication, and collaboration skills, with a clear sense of ownership; experience with AI performance monitoring, vector database challenges at scale, LLM latency analysis, and AI tools such as MCP and agents is a plus