CareerRiver

AI Implementation Quality Analyst

Granicus, LLC · Remote

📍 Remote, UNAVAILABLE💰 $80,000-$105,700via icimsPosted 2026-06-10
Apply on company site ↗
CareerRiver pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Granicus, LLC.
The Company Serving the People Who Serve the People Granicus is driven by the excitement of building, implementing, and maintaining technology that is transforming the Govtech industry by bringing governments and its constituents together. We are on a mission to support our customers with meeting the needs of their communities and implementing our technology in ways that are equitable and inclusive. Granicus has consistently appeared on the GovTech 100 list over the past 5 years and has been recognized as the best companies to work on BuiltIn. Over the last 25 years, we have served 5,500 federal, state, and local government agencies and more than 300 million citizen subscribers power an unmatched Subscriber Network that use our digital solutions to make the world a better place. With comprehensive cloud-based solutions for communications, government website design, meeting and agenda management software, records management, and digital services, Granicus empowers stronger relationships between government and residents across the U.S., U.K., Australia, New Zealand, and Canada. By simplifying interactions with residents, while disseminating critical information, Granicus brings governments closer to the people they serve—driving meaningful change for communities around the globe. Want to know more? See more of what we do here. Job Summary The AI Quality & Evaluation Analyst is responsible for assessing the quality, correctness, completeness, and safety of AI‑generated responses across defined use cases. This role combines hands‑on human review with structured, rubric‑based evaluation incorporating automation to ensure AI systems meet documented standards before customer implementations go live.  As part of the implementation team, this role is client-facing and serves as a bridge between client expectations and system behavior, translating real-world use cases, domain context, and risk tolerance into measurable evaluation criteria and actionable feedback for product and engineering teams. The role focuses on what the AI says and does, not on model training or infrastructure performance. The analyst serves as a human quality gate, ensuring outputs are accurate, appropriate for the audience, and aligned with policy and governance requirements.  This role works closely with product and domain experts to translate real‑world expectations into measurable evaluation criteria and repeatable test artifacts.   What Your Impact Will Look Like Partner directly with clients during implementation to understand use cases, success criteria, and risk tolerance Translate client requirements into evaluation frameworks, prompt strategies, and test coverage Act as the quality liaison between client, product, and engineering to ensure alignment pre- and post-launch Review and score AI responses using standard rubrics  Validate against sources/ground truth; flag hallucinations, omissions, and other risks  Document findings and calibrate with peers to keep scoring consistent  Build and maintain prompt banks and golden sets (expected results)  Expand coverage for edge cases, high-risk scenarios, and real user language  Track regression/drift and feed dashboards and quality reports  Triage internal and client feedback; synthesize themes across deployments Identify systemic risks and escalate high-impact findings with clear, client-relevant context Partner with product/engineering to validate and verify fixes, ensuring alignment with client expectations You Will Love This Job If You Have Strong analytical judgment and attention to detail  Excellent written communication skills, including the ability to explain reasoning clearly  Experience reviewing, auditing, or evaluating structured outputs such as content, decisions, or recommendations  Comfort applying detailed guidelines and rubrics consistently at scale  Familiarity with large language models and common failure modes such as hallucinations, overgeneralization, or unsafe responses  Ability to work in spreadsheets, evaluation tools, or annotation platforms   Experience with AI evaluation, data annotation, QA, trust and safety, or policy review  Exposure to human‑in‑the‑loop workflows, or benchmarking processes  Domain expertise in a regulated or high‑risk field such as government, education, healthcare, or legal services  Experience contributing to test suites, evaluation dashboards, or quality reporting  Comfort working cross‑functionally with product and engineering teams Experience in client-facing roles such as implementation, consulting, or solution delivery in a SaaS or regulated environment Ability to translate between business/user needs and technical system behavior AI responses consistently meet defined standards for accuracy, completeness, and safety in real client use cases Evaluation results are reproducible and trusted across teams and with clients Client requirements and risk tolerances are clearly translated into evaluation criteria and test coverage Quality issues are identified early, clearly documented, and actionable  Test coverage expands over time to reflect real user behavior and risk  Stakeholders can confidently use evaluation outputs to make release and governance decisions   AI outputs are consistently accurate, complete, and safe in real client use cases, with quality standards that are clearly defined and trusted across teams. Client expectations and risk tolerances are effectively translated into repeatable evaluation frameworks and test coverage. Issues are identified early, communicated clearly, and resolved quickly with engineering. Evaluation outputs are reliable and actionable, enabling confident release decisions and building trust with both internal stakeholders and clients. A strong instinct for what “good” looks like and the discipline to define and measure it. You enjoy digging into ambiguous outputs, spotting subtle issues, and bac

More Remote jobs

Remote jobs · Browse all locations