Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
Outlier is hiring worldwide — AI Trainers, Coding Experts & Writing EvaluatorsFreelance • Fully remote • $15–$60/hr • Work when you wantBrowse Outlier roles on MMagic.ai →
W

Senior Prompt Engineer (Remote — Worldwide)

Writer$150k – $260kRemote (Global)Posted 1w ago
Prompt EngineeringPythonPalmyra LLMNLPRAG OptimizationModel EvaluationLangChainGenerative AI SafetySystem Prompt Design

About the role

As a Senior Prompt Engineer at Writer, you will bridge the gap between enterprise raw LLM capabilities and production-grade applications for the world's leading brands. You will be responsible for architecting sophisticated prompt systems and evaluation frameworks that power our Palmyra family of models and enterprise-grade RAG pipelines. Your work will directly impact how Fortune 500 companies deploy AI that is accurate, brand-aligned, and secure, ensuring our platform delivers reliable outputs at scale.

Responsibilities

- Design, test, and iterate on complex system prompts for Writer's Palmyra models to support diverse enterprise use cases. - Develop and maintain comprehensive evaluation suites to measure model performance, reliability, and safety across different industry verticals. - Research and implement cutting-edge prompting techniques such as Chain-of-Thought, ReAct, and Tree-of-Thought within our platform. - Collaborate with the NLP research team to provide feedback for model fine-tuning based on observed prompt performance. - Optimize prompts for latency and cost without compromising the quality or accuracy of the generated output. - Build "Golden Datasets" and benchmark suites to validate platform updates and prevent regression in model responses. - Advocate for prompt engineering best practices across the engineering and product organizations.

Requirements

- 4+ years of experience working with Large Language Models, specifically in prompt engineering and fine-tuning contexts. - Demonstrated experience building complex, multi-step prompt chains and agentic workflows. - Strong background in building automated evaluation datasets (LLM-as-a-judge) and scoring rubrics. - Proficiency in Python and experience with orchestration frameworks like LangChain or LlamaIndex. - Deep understanding of RAG (Retrieval-Augmented Generation) architectures and how to optimize prompts for context retrieval. - Experience with enterprise constraints including data privacy, brand voice consistency, and hallucinations. - Excellent communication skills to translate business requirements into technical prompt specifications.

Benefits

Equity, benefits, remote.
Disclaimer: MMagic.ai connects talented people with AI companies around the world. While we work hard to feature quality opportunities, we don't independently verify employers, candidates, salaries, or hiring outcomes. We encourage you to research each opportunity and company before applying or making an offer.

Similar roles