AI Implementation Quality Analyst
Granicus, LLC · Remote
📍 Remote, UNAVAILABLE💰 $80,000-$105,700via icimsPosted 2026-06-10
Apply on company site ↗
CareerRiver pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Granicus, LLC.
The Company
Serving the People Who Serve the People
Granicus is driven by the excitement of building, implementing, and maintaining technology that is transforming the Govtech industry by bringing governments and its constituents together. We are on a mission to support our customers with meeting the needs of their communities and implementing our technology in ways that are equitable and inclusive. Granicus has consistently appeared on the GovTech 100 list over the past 5 years and has been recognized as the best companies to work on BuiltIn.
Over the last 25 years, we have served 5,500 federal, state, and local government agencies and more than 300 million citizen subscribers power an unmatched Subscriber Network that use our digital solutions to make the world a better place. With comprehensive cloud-based solutions for communications, government website design, meeting and agenda management software, records management, and digital services, Granicus empowers stronger relationships between government and residents across the U.S., U.K., Australia, New Zealand, and Canada. By simplifying interactions with residents, while disseminating critical information, Granicus brings governments closer to the people they serve—driving meaningful change for communities around the globe.
Want to know more? See more of what we do here.
Job Summary
The AI Quality & Evaluation Analyst is responsible for assessing the quality, correctness, completeness, and safety of AI‑generated responses across defined use cases. This role combines hands‑on human review with structured, rubric‑based evaluation incorporating automation to ensure AI systems meet documented standards before customer implementations go live.
As part of the implementation team, this role is client-facing and serves as a bridge between client expectations and system behavior, translating real-world use cases, domain context, and risk tolerance into measurable evaluation criteria and actionable feedback for product and engineering teams.
The role focuses on what the AI says and does, not on model training or infrastructure performance. The analyst serves as a human quality gate, ensuring outputs are accurate, appropriate for the audience, and aligned with policy and governance requirements.
This role works closely with product and domain experts to translate real‑world expectations into measurable evaluation criteria and repeatable test artifacts.
What Your Impact Will Look Like
Partner directly with clients during implementation to understand use cases, success criteria, and risk tolerance
Translate client requirements into evaluation frameworks, prompt strategies, and test coverage
Act as the quality liaison between client, product, and engineering to ensure alignment pre- and post-launch
Review and score AI responses using standard rubrics
Validate against sources/ground truth; flag hallucinations, omissions, and other risks
Document findings and calibrate with peers to keep scoring consistent
Build and maintain prompt banks and golden sets (expected results)
Expand coverage for edge cases, high-risk scenarios, and real user language
Track regression/drift and feed dashboards and quality reports
Triage internal and client feedback; synthesize themes across deployments
Identify systemic risks and escalate high-impact findings with clear, client-relevant context
Partner with product/engineering to validate and verify fixes, ensuring alignment with client expectations
You Will Love This Job If You Have
Strong analytical judgment and attention to detail
Excellent written communication skills, including the ability to explain reasoning clearly
Experience reviewing, auditing, or evaluating structured outputs such as content, decisions, or recommendations
Comfort applying detailed guidelines and rubrics consistently at scale
Familiarity with large language models and common failure modes such as hallucinations, overgeneralization, or unsafe responses
Ability to work in spreadsheets, evaluation tools, or annotation platforms
Experience with AI evaluation, data annotation, QA, trust and safety, or policy review
Exposure to human‑in‑the‑loop workflows, or benchmarking processes
Domain expertise in a regulated or high‑risk field such as government, education, healthcare, or legal services
Experience contributing to test suites, evaluation dashboards, or quality reporting
Comfort working cross‑functionally with product and engineering teams
Experience in client-facing roles such as implementation, consulting, or solution delivery in a SaaS or regulated environment
Ability to translate between business/user needs and technical system behavior
AI responses consistently meet defined standards for accuracy, completeness, and safety in real client use cases
Evaluation results are reproducible and trusted across teams and with clients
Client requirements and risk tolerances are clearly translated into evaluation criteria and test coverage
Quality issues are identified early, clearly documented, and actionable
Test coverage expands over time to reflect real user behavior and risk
Stakeholders can confidently use evaluation outputs to make release and governance decisions
AI outputs are consistently accurate, complete, and safe in real client use cases, with quality standards that are clearly defined and trusted across teams. Client expectations and risk tolerances are effectively translated into repeatable evaluation frameworks and test coverage. Issues are identified early, communicated clearly, and resolved quickly with engineering. Evaluation outputs are reliable and actionable, enabling confident release decisions and building trust with both internal stakeholders and clients.
A strong instinct for what “good” looks like and the discipline to define and measure it. You enjoy digging into ambiguous outputs, spotting subtle issues, and bac
More Remote jobs
Remote jobs · Browse all locations