CareerRiver

AI Enablement & Governance– AI Quality & Evaluation Lead

Alight · Remote

📍 US-IL-Illinois-Virtualvia workday
Apply on company site ↗
CareerRiver pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Alight.
Our story At Alight, we believe a company’s success starts with its people. At our core, we Champion People, help our colleagues Grow with Purpose and true to our name we encourage colleagues to “Be Alight.”  We are passionate about connecting purpose with impact.  Alight empowers clients to build a healthier and more financially secure workforce by unifying the benefits ecosystem across health, wealth, wellbeing, navigation, and absence management. Our Benefits With a comprehensive total rewards package, Alight offers programs and plans that support your mind, body, wallet, and life. Benefits include health, dental and vision coverages starting Day One. Additionally, Alight colleagues enjoy wellbeing programs, retirement plans with contribution matching, generous time off, parental leave, continuing education, and career growth opportunities – all within a thriving global organization. Flexible Working So that you can be your best at work and home, we consider flexible working arrangements wherever possible. Alight has been a leader in the flexible workspace and “Top 100 Company for Remote Jobs” 6 years in a row. Great Place to Work Thanks to the work of every colleague, Alight has received multiple awards of recognition including “Great Place to Work” for the past 7 years and Fortune’s “Best Companies to Work For.” To learn more about our company culture and awards Click Here. If you, Champion People, seek to Grow with Purpose, and embody the meaning of Be Alight – We invite you to join our team!  Learn more at careers.alight.com .  The Role The AI Quality & Evaluation lead enables responsible and scalable AI adoption by defining technical quality standards, evaluation framework and control requirements across the AI lifecycle. As a function owner of AI quality and robustness standards, the role translates enterprise trust principles and statistical rigor into practical evaluation standards and enforceable technical controls.  The role partners closely with AI Engineering, Data Scientists, QA CoE, and Product teams to define quality and performance requirements that are embedded by design, ensuring AI solutions—including RAG systems and Agents—are accurate, grounded, and aligned with enterprise risk and trust expectations. Responsibilities Quality-by-Design Partnership Partnering directly with AI Engineers, Application Developers and Data Scientists during the design phase to define technical quality acceptance criteria and fit-for-use requirements. Embedding quality considerations into model and system architecture from the onset, specifically for complex patterns like RAG and autonomous Agents. Defining golden truth requirements and evaluation dataset standards; partner with Data Science teams to ensure datasets reflect production-level complexity. Defining quality and evaluation expectations for third-party AI systems and vendor-supplied models, ensuring consistent governance standards regardless of model origin. Technical Evaluation & Metric Engineering Designing and maintaining structured evaluation framework that assesses AI system against defined quality bars (e.g. Goodness-of-Fit, Calibration, Stability). Developing automated metrics for Generative AI performance including Groundedness (Hallucination detection), Faithfulness, Completeness, and other domain-relevant metrics. Defining and operationalize fairness and bias evaluation criteria, including demographic parity assessments and disparate impact testing for client-facing AI systems. Calibrating evaluation thresholds and monitoring cadence to AI risk tier, ensuring proportionate controls without over-engineering lower-risk use cases.  Technical Control & Monitoring Identifying and document technical AI governance controls that enable automated compliance with performance and risk obligations. Establishing drift and ongoing monitoring requirements, defining statistical triggers for feature and concept drift that necessitate model intervention. Developing clear control statements that articulate the expected evidence artifacts (e.g. test results, model cards) required go/no-go decisions. Governance & Evidence Enablement Providing objective, data-driven evaluation outputs that support AI governance reviews and risk classification. Translating governance expectations into clear, testable quality criteria that engineering teams can apply consistently within their CI/CD pipelines. Maintaining authoritative documentation of AI controls to support audit, regulatory review, and internal assurance activities. Requirements Technical Depth: 5–8+ years of experience in Data Science, ML Engineering, or AI Quality, with a focus on evaluation and statistical validation. System Design: Practical experience partnering with engineers to design RAG, LLM-based Agents, or traditional ML pipelines. Analytical Skills: Expert-level Python (Pandas, Scikit-learn) and experience with evaluation frameworks (e.g., RAGAS, TruLens, or MLflow). Governance Mindset: Demonstrated ability to translate abstract trust concepts into mathematical metrics and enforceable technical controls. Stakeholder Influence: Demonstrated ability to work without direct authority, driving quality adoption across engineering and product teams through enablement.  Communication: Ability to bridge the gap between high-level governance policy and low-level code implementation. Bachelor’s degree in a technical field (e.g., Computer Science, Computer Systems Design) or equivalent professional experience Application and Interview By applying for a position with Alight, you understand that, should you be made an offer, it will be contingent on your undergoing and successfully completing a background check consistent with Alight’s employment policies. Background checks may include some or all the following based on the nature of the position: SSN/SIN validation, education verification, employment verification, and crimina

More Remote jobs

Remote jobs · Browse all locations