I will test your ai chatbot or agent for hallucinations and UX issues
SaaS AI and Indie Game QA Tester and Narrative Designer
Nivel 1
Ha cumplido determinados criterios de rendimiento y muestra un gran potencial en la plataforma.
Acerca de este Servicio
AI assistants often work during controlled demos but fail when real users provide incomplete information, vague questions, conflicting instructions, unexpected follow-ups, or requests outside the ideal path.
I will manually evaluate your AI chatbot, RAG assistant, customer support bot, or tool-using AI agent using structured, real-world test scenarios.
Testing can cover:
- Response accuracy and relevance
- Unsupported or hallucinated claims
- Instruction following
- Multi-turn consistency
- Missing or ambiguous information
- RAG grounding against supplied documents
- Tone and customer experience
- Refusal and escalation behaviour
- Basic prompt-injection behaviour checks
- Broken UI and tool workflows
You can receive:
- Prompt and response evidence
- Expected behaviour
- Failure description
- Severity and business impact
- Screenshots
- Recommended improvements
- A reusable regression prompt set
This is a functional and quality evaluation. It is not a formal security penetration test, compliance audit, model fine-tuning service, or guarantee that every possible AI failure will be discovered.
Please contact me before ordering if your agent uses private data, multiple external tools, regulated information, or complex
Aplicación de prueba:
Software
Tecnología de desarrollo:
C/C++
•
Flutter
•
Java
•
JavaScript
•
Python
Dispositivo:
PC
