
on a mission:
Helping people train, learn, and collaborate across languages — no matter where they are (or which reality they're in).
“A bit of a nightmare,” they say — and I’m instantly intrigued. A casual chat about a user study turns into me joining the project at full speed — right near the finish line, of course. We scrap the old plan, start from zero, and race the clock, and turn chaos into clarity.
00 Overview
Teaching AI to Train Us
About the application
AIXTRA is a VR training application developed within the EU-funded VOXReality initiative. It integrates real-time AI translation to overcome language barriers and an AI-based virtual training partner to simulate human-to-human collaboration.
Info
Project type: UX/UI, “Add a Feature”
Medium: Mobile application
Role: UX/UI Designer, Researcher
Team: 11
Problem
AI systems are increasingly deployed in training environments, but concerns remain about their ability to correctly recognise user intent, provide accurate translations, and build user trust. This study explored these issues through two targeted experiments.
Outcome
User Study:
Total participants: 36
Test sessions: total 24 sessions across 6 days.
01 problem
Does AI Really Get What You Mean — and Help You Learn?
When I joined theproject, the goals were undefined and the requirements unclear. The VR training app has never been tested by anyone other than the developers, and it was built intuitively, without a foundation in UX or UI principles. It’s a tough nut to crack — but also a great opportunity to apply my skills and, in the long run, help the developers shift their perspective for future projects.
It quickly became clear that before testing could begin, we needed to take a step back — define who the users really were, and what problems the AI system was meant to solve.
The app’s core functionalities:
AI-Based Virtual Training Partner
Real-Time AI Translation
The project’s limitations:
1
The app has been already developed without clear user test requirements in mind.
2
No possibilities to test with a control group.
3
Arbitrary requirement to perform a quantitative study with 30 participants.
4
Limited time & resources.
Opportunity:
The AIXTRA app was build to provide a practical use case for presenting improved interaction between humans and AI in XR environments. AIXTRA aims to enhance VR training by overcoming language barriers and providing real-time guidance through a pretrained VOXReality AI-driven virtual assistant and translation system.
GoalS:
The tests aim to determine whether intent recognition, live translation, and AI support can improve communication, user confidence, and overall training effectiveness.
Research objectives:
Evaluate the functional accuracy of intent recognition systems.
Does it understand and execute user commands and requests?
Asses the usefulness and quality of real-time AI translation features.
Does it contribute positively to team-based interactions among multilingual users?
Examine the role of AI assistance in enhancing task and training performance, as well as user experience.
Do users perceive the AI assistant as helpful and supportive in achieving training objectives?
Verify the inclusivity and accessibility benefits of integrating real-time translation in multi-cultural training contexts.
Does it reduce communication barriers?
How do different language patterns affect the performance?
Analyse user attitudes towards the integration of AI systems (ethics, accuracy, functionality, data protection).
Investigate how the presence of an AI-based Training Partner affects the user autonomy.
Overall assumptions:
02 Solution
How We Measured AI’s Smarts in VR Training
divide and conquer
1
Test: AI-based Training Partner
What we mainly test: Intent Recognition
How: Users participate in a single-person training scenario
2
Test: AI real-time translation
What we mainly test: Translation Efficiency and Impact
How: Users participate in a multi-person training scenario in groups of 2 or 3.
Before I joined, a preliminary survey of over 30 mostly open-ended questions had been created. I began by refining it by analysing each question to ensure clarity, relevance, and measurable value.
My focus was on making questions more direct, reducing bias, and selecting appropriate response types to enable meaningful comparison and reliable analysis of user feedback.
2
language versions - English & German
2
parts: a pre-test questionnaire and a post-test evaluation
aim
to understand how AI intent recognition and translation affect user confidence, communication, inclusivity, and willingness to adopt AI-assisted VR training systems.
pre-test section
gathers demographic information, language background, familiarity with AI and XR technologies, and initial expectations or concerns about AI in training contexts.
post-test section
participants rate aspects such as accuracy, response time, trust, human-likeness, privacy concerns, and perceived usefulness. Open-ended and multiple-choice questions allow participants to share insights about missing features, emotional engagement, and comfort with data processing trade-offs.
While my colleagues recruited participants across five different channels, I focused on preparing clear, user-friendly consent forms and participant handouts with user stories to help them understand the training scenarios they were to complete.
THe Reactor Cleaning VR training
1
Single User Scenario:Users complete a short tutorial on movement and tool use before starting training. As the Cleaner, they follow written and spoken instructions from the AI Assistant to safely shut down, open, and clean the reactor.
2
Multi-User Scenario:
Three roles work together - Operator, Cleaner, and Assistant.The Operator monitors the dashboard, gives verbal instructions, and confirms task progress.The Cleaner performs the reactor cleaning, following the Operator’s directions and communicating through a wrist walkie-talkie.The Assistant supports the Cleaner by handing tools and coordinating with both teammates.

01
Single User story

02
Multi-User sample story for Operator
Can we clean the reactor?
1 Trying outthe app
Users participated in 45 minute sessions, completing the tasks in the VR scenario and filling out the surveys. Live observations and follow-up questions proved useful.
What we prepared:

2 Application Audit
Additionally, I made many observations during the tests and identified points of improvement for designing training applications.

03 Results
Reality Check: When AIXTRA Met Real Users
Methodology:
36 participants
16 users tested the single-user scenario.
24 users tested the multi-user scenario in groups of 2 or 3.Diverse linguistic backgrounds, including native and non-native speakers, different age groups, and various experiences with AI and XR technology.
Procedure
Participants completed structured tasks in the VR scenario.
Quantitative error metrics were collected through Unity logs. Quantitative and qualitative user feedback was collected via a survey to capture real-time performance and user experience.
Data collection
Survey feedback (quantitative & qualitative), direct observation.
Used Likert scales and multiple-choice questions to evaluate.







Key findings:
Intent Recognition:
Overall functional, with correct interpretations in most cases, but less accurate with ambiguous or colloquial phrasing.
Translation Accuracy:
High in simple contexts; reduced accuracy in domain-specific or idiomatic expressions. Occasional mistranslations altered the meaning.
User Experience:
Participants valued speed and accessibility. However, lack of nuance in translations and occasional contextual mismatches caused confusion.
User Concerns :
While users recognise the practical benefits of AI-assisted training systems, a significant portion voice concerns about data privacy, potential algorithmic bias, and the reliability of AI responses.
User study evaluation and improvement points:
Misalignment Between Users & Target Audience
Test participants lacked the domain knowledge expected by the app, leading to confusion.
What I suggest
Ensure test users are either pre-briefed or provided with clear, accessible materials (visual aids, simple instructions, minimal jargon) before using the application.
AI Assistant’s Lack of Contextual Awareness
Even when users sought help, the AI assistant could not assess the user’s current situation or reason about what support was needed.
Assistance was only effective when highly specific questions were asked.
What I suggest
Provide clear definitions, images, or tutorial content for key equipment.
Lack of Clear Feedback & Completion Cues
Users were unsure whether tasks were correctly completed.
They expected the assistant to monitor their progress like a human assistant might, but the assistant only reacts to explicit queries.
The application's interface misleadingly implies that the assistant is actively involved, fostering false expectations about its capabilities.
What I suggest
Set realistic expectations about the assistant’s role and provide more explicit task progression indicators and completion confirmations.
Breakdown in Multi-User Scenarios
Operators unfamiliar with procedures caused confusion, amplified by weak translations.
What I suggest
For multi-user tests, only assign the operator role to someone already familiar with the procedure. Let the cleaner and assistant rely on this knowledgeable operator to simulate the intended dynamic.
Maintaining the Learning Value of the VR Experience
Although users are operating in a safe, simulated environment, the goal is still effective learning.
One user mentioned that due to frequent back-and-forth with the assistant and translation issues, she wasn’t confident she'd retained the full process after completing the training.
Users often didn’t finish tasks or learn anything due to accumulated friction - from confusing tools, unclear instructions, or non-intuitive controls. These obstacles discouraged exploration and led to avoidance of the assistant or key functionality.
My insight
Cognitive overload, caused by translation friction and repeated clarification, can erode learning outcomes -even if the task is eventually completed.
the Role and Perception of the AI Assistant
Users were unaware of the AI assistant’s capabilities, leading to underuse of features like object highlighting. Misaligned expectations also emerged, as many assumed the assistant would act like a human colleague.
My recommendation
Explicitly communicate the assistant's capabilities.
Make key features more obvious (e.g. blinking highlights, greying out irrelevant elements).
Consider offering onboarding prompts and expanding the tutorial demonstrating assistant interactions.
A few notes on translation feature testing:
Evaluating Translation Within the Training Context
Users remained alert to translation errors, questioning odd or unsafe phrases like “breathe inside the reactor.” This critical attitude shows strong engagement and situational awareness — a positive indicator that translation issues did not compromise understanding or safety.
My insight
Users are capable of detecting mistranslations. However, subtle mistranslations may go unnoticed, which could affect task accuracy or safety.
The Role of the Operator in Translation Success
When operators were unsure of the procedure, their instructions became lengthy, hesitant, incomplete or fragmented. This ambiguity led to lower translation quality and confusion downstream.
Principle
Clear source communication = better translation quality. Operators should be well-trained or scripted to ensure consistent, translatable instructions.
Translation Shortening and Information Loss
The translation system often condensed long, rambling operator responses, which improved clarity.
However, there is a risk that key information may be omitted during this shortening process.
My recommendation
Implement post-session checks to evaluate whether essential information is preserved in translated responses.

Users become frustrated when basic interactions don’t match real-world expectations. For example, in the wrench tutorial, they turned the tool incorrectly because the interface didn’t mimic real-world behaviour. Interfaces should align with real-world expectations and muscle memory.
In physical training, a mentor guides learners step by step, but in VR this guidance must be replaced with tools like manuals, diagrams, or visual cues. Current implementations, such as hovering instruction boards, are often cluttered and poorly positioned. Provide just-in-time instructions placed at eye level or within reach to make guidance intuitive and easily accessible.
Thanks for training with us!Care for projects with more UI focus?

on a mission:
Helping people train, learn, and collaborate across languages — no matter where they are (or which reality they're in).
“A bit of a nightmare,” they say — and I’m instantly intrigued. A casual chat about a user study turns into me joining the project at full speed — right near the finish line, of course. We scrap the old plan, start from zero, and race the clock, and turn chaos into clarity.
00
Overview
Teaching AI to Train Us
About the application
AIXTRA is a VR training application developed within the EU-funded VOXReality initiative. It integrates real-time AI translation to overcome language barriers and an AI-based virtual training partner to simulate human-to-human collaboration.
Info
Project type: VR Application audit and testing
Medium: VR Training
Role: UX Researcher
Team: 11
Problem
AI systems are increasingly deployed in training environments, but concerns remain about their ability to correctly recognise user intent, provide accurate translations, and build user trust. This study explored these issues through two targeted experiments.
Outcome
User Study:
Total participants: 36
Test sessions: total 24 sessions across 6 days.
01
problem
Does AI Really Get What You Mean — and Help You Learn?
When I joined theproject, the goals were undefined and the requirements unclear. The VR training app has never been tested by anyone other than the developers, and it was built intuitively, without a foundation in UX or UI principles. It’s a tough nut to crack — but also a great opportunity to apply my skills and, in the long run, help the developers shift their perspective for future projects.
It quickly became clear that before testing could begin, we needed to take a step back — define who the users really were, and what problems the AI system was meant to solve.
The app’s core functionalities:
AI-Based Virtual Training Partner
Real-Time AI Translation
The project’s limitations:
1
The app has been already developed without clear user test requirements in mind.
2
No possibilities to test with a control group.
3
Arbitrary requirement to perform a quantitative study with 30 participants.
4
Limited time & resources.
Opportunity:
The AIXTRA app was build to provide a practical use case for presenting improved interaction between humans and AI in XR environments. AIXTRA aims to enhance VR training by overcoming language barriers and providing real-time guidance through a pretrained VOXReality AI-driven virtual assistant and translation system.
GoalS:
The tests aim to determine whether intent recognition, live translation, and AI support can improve communication, user confidence, and overall training effectiveness.
Research objectives:
Evaluate the functional accuracy of intent recognition systems.
Does it understand and execute user commands and requests?
Asses the usefulness and quality of real-time AI translation features.
Does it contribute positively to team-based interactions among multilingual users?
Examine the role of AI assistance in enhancing task and training performance, as well as user experience.
Do users perceive the AI assistant as helpful and supportive in achieving training objectives?
Verify the inclusivity and accessibility benefits of integrating real-time translation in multi-cultural training contexts.
Does it reduce communication barriers?
How do different language patterns affect the performance?
Analyse user attitudes towards the integration of AI systems (ethics, accuracy, functionality, data protection).
Investigate how the presence of an AI-based Training Partner affects the user autonomy.
Overall assumptions:
02
Solution
How We Measured AI’s Smarts in VR Training
divide and conquer
1
Test: AI-based Training Partner
What we mainly test: Intent Recognition
How: Users participate in a single-person training scenario
2
Test: AI real-time translation
What we mainly test: Translation Efficiency and Impact
How: Users participate in a multi-person training scenario in groups of 2 or 3.
Before I joined, a preliminary survey of over 30 mostly open-ended questions had been created. I began by refining it by analysing each question to ensure clarity, relevance, and measurable value.
My focus was on making questions more direct, reducing bias, and selecting appropriate response types to enable meaningful comparison and reliable analysis of user feedback.
2
language versions - English & German
2
parts: a pre-test questionnaire and a post-test evaluation
aim
to understand how AI intent recognition and translation affect user confidence, communication, inclusivity, and willingness to adopt AI-assisted VR training systems.
pre-test section
gathers demographic information, language background, familiarity with AI and XR technologies, and initial expectations or concerns about AI in training contexts.
post-test section
participants rate aspects such as accuracy, response time, trust, human-likeness, privacy concerns, and perceived usefulness. Open-ended and multiple-choice questions allow participants to share insights about missing features, emotional engagement, and comfort with data processing trade-offs.
While my colleagues recruited participants across five different channels, I focused on preparing clear, user-friendly consent forms and participant handouts with user stories to help them understand the training scenarios they were to complete.
THe Reactor Cleaning VR training
1
Single User Scenario:Users complete a short tutorial on movement and tool use before starting training. As the Cleaner, they follow written and spoken instructions from the AI Assistant to safely shut down, open, and clean the reactor.
2
Multi-User Scenario:
Three roles work together - Operator, Cleaner, and Assistant.The Operator monitors the dashboard, gives verbal instructions, and confirms task progress.The Cleaner performs the reactor cleaning, following the Operator’s directions and communicating through a wrist walkie-talkie.The Assistant supports the Cleaner by handing tools and coordinating with both teammates.

01
Single User story

02
Multi-User sample story for Operator
Can we clean the reactor?
1 Trying outthe app
Users participated in 45 minute sessions, completing the tasks in the VR scenario and filling out the surveys. Live observations and follow-up questions proved useful.
What we prepared:


2 Application Audit
Additionally, I made many observations during the tests and identified points of improvement for designing training applications.
03
Results
Reality Check: When AIXTRA Met Real Users
Methodology:
36 participants
16 users tested the single-user scenario.
24 users tested the multi-user scenario in groups of 2 or 3.Diverse linguistic backgrounds, including native and non-native speakers, different age groups, and various experiences with AI and XR technology.
Procedure
Participants completed structured tasks in the VR scenario.
Quantitative error metrics were collected through Unity logs. Quantitative and qualitative user feedback was collected via a survey to capture real-time performance and user experience.
Data collection
Survey feedback (quantitative & qualitative), direct observation.
Used Likert scales and multiple-choice questions to evaluate.







Key findings:
Intent Recognition:
Overall functional, with correct interpretations in most cases, but less accurate with ambiguous or colloquial phrasing.
Translation Accuracy:
High in simple contexts; reduced accuracy in domain-specific or idiomatic expressions. Occasional mistranslations altered the meaning.
User Experience:
Participants valued speed and accessibility. However, lack of nuance in translations and occasional contextual mismatches caused confusion.
User Concerns :
While users recognise the practical benefits of AI-assisted training systems, a significant portion voice concerns about data privacy, potential algorithmic bias, and the reliability of AI responses.
User study evaluation and improvement points:
Misalignment Between Users & Target Audience
Test participants lacked the domain knowledge expected by the app, leading to confusion.
What I suggest
Ensure test users are either pre-briefed or provided with clear, accessible materials (visual aids, simple instructions, minimal jargon) before using the application.
AI Assistant’s Lack of Contextual Awareness
Even when users sought help, the AI assistant could not assess the user’s current situation or reason about what support was needed.
Assistance was only effective when highly specific questions were asked.
What I suggest
Provide clear definitions, images, or tutorial content for key equipment.
Lack of Clear Feedback & Completion Cues
Users were unsure whether tasks were correctly completed.
They expected the assistant to monitor their progress like a human assistant might, but the assistant only reacts to explicit queries.
The application's interface misleadingly implies that the assistant is actively involved, fostering false expectations about its capabilities.
What I suggest
Set realistic expectations about the assistant’s role and provide more explicit task progression indicators and completion confirmations.
Breakdown in Multi-User Scenarios
Operators unfamiliar with procedures caused confusion, amplified by weak translations.
What I suggest
For multi-user tests, only assign the operator role to someone already familiar with the procedure. Let the cleaner and assistant rely on this knowledgeable operator to simulate the intended dynamic.
Maintaining the Learning Value of the VR Experience
Although users are operating in a safe, simulated environment, the goal is still effective learning.
One user mentioned that due to frequent back-and-forth with the assistant and translation issues, she wasn’t confident she'd retained the full process after completing the training.
Users often didn’t finish tasks or learn anything due to accumulated friction - from confusing tools, unclear instructions, or non-intuitive controls. These obstacles discouraged exploration and led to avoidance of the assistant or key functionality.
My insight
Cognitive overload, caused by translation friction and repeated clarification, can erode learning outcomes -even if the task is eventually completed.
the Role and Perception of the AI Assistant
Users were unaware of the AI assistant’s capabilities, leading to underuse of features like object highlighting. Misaligned expectations also emerged, as many assumed the assistant would act like a human colleague.
My recommendation
Explicitly communicate the assistant's capabilities.
Make key features more obvious (e.g. blinking highlights, greying out irrelevant elements).
Consider offering onboarding prompts and expanding the tutorial demonstrating assistant interactions.
A few notes on translation feature testing:
Evaluating Translation Within the Training Context
Users remained alert to translation errors, questioning odd or unsafe phrases like “breathe inside the reactor.” This critical attitude shows strong engagement and situational awareness — a positive indicator that translation issues did not compromise understanding or safety.
My insight
Users are capable of detecting mistranslations. However, subtle mistranslations may go unnoticed, which could affect task accuracy or safety.
The Role of the Operator in Translation Success
When operators were unsure of the procedure, their instructions became lengthy, hesitant, incomplete or fragmented. This ambiguity led to lower translation quality and confusion downstream.
Principle
Clear source communication = better translation quality. Operators should be well-trained or scripted to ensure consistent, translatable instructions.
Translation Shortening and Information Loss
The translation system often condensed long, rambling operator responses, which improved clarity.
However, there is a risk that key information may be omitted during this shortening process.
My recommendation
Implement post-session checks to evaluate whether essential information is preserved in translated responses.

Users become frustrated when basic interactions don’t match real-world expectations. For example, in the wrench tutorial, they turned the tool incorrectly because the interface didn’t mimic real-world behaviour. Interfaces should align with real-world expectations and muscle memory.
In physical training, a mentor guides learners step by step, but in VR this guidance must be replaced with tools like manuals, diagrams, or visual cues. Current implementations, such as hovering instruction boards, are often cluttered and poorly positioned. Provide just-in-time instructions placed at eye level or within reach to make guidance intuitive and easily accessible.
Thanks for training with us!Care for projects with more UI focus?

on a mission:
Helping people train, learn, and collaborate across languages — no matter where they are (or which reality they're in).
“A bit of a nightmare,” they say — and I’m instantly intrigued. A casual chat about a user study turns into me joining the project at full speed — right near the finish line, of course. We scrap the old plan, start from zero, and race the clock, and turn chaos into clarity.
00
Overview
Teaching AI to Train Us
About the application
AIXTRA is a VR training application developed within the EU-funded VOXReality initiative. It integrates real-time AI translation to overcome language barriers and an AI-based virtual training partner to simulate human-to-human collaboration.
Info
Project type: VR Application audit and testing
Medium: VR Training
Role: UX Researcher
Team: 11
Problem
AI systems are increasingly deployed in training environments, but concerns remain about their ability to correctly recognise user intent, provide accurate translations, and build user trust. This study explored these issues through two targeted experiments.
Outcome
User Study:
Total participants: 36
Test sessions: total 24 sessions across 6 days.
01
problem
Does AI Really Get What You Mean — and Help You Learn?
When I joined theproject, the goals were undefined and the requirements unclear. The VR training app has never been tested by anyone other than the developers, and it was built intuitively, without a foundation in UX or UI principles. It’s a tough nut to crack — but also a great opportunity to apply my skills and, in the long run, help the developers shift their perspective for future projects.
It quickly became clear that before testing could begin, we needed to take a step back — define who the users really were, and what problems VOXReality's AI technology was meant to solve.
The app’s core functionalities:
AI-Based Virtual Training Partner
Real-Time AI Translation
The project’s limitations:
1
The app has been already developed without clear user test requirements in mind.
2
No possibilities to test with a control group.
3
Arbitrary requirement to perform a quantitative study with 30 participants.
4
Limited time & resources.
Opportunity:
The AIXTRA app was build to provide a practical use case for presenting improved interaction between humans and AI in XR environments. AIXTRA aims to enhance VR training by overcoming language barriers and providing real-time guidance through a pretrained VOXReality AI-driven virtual assistant and translation system.
GoalS:
The tests aim to determine whether intent recognition, live translation, and AI support can improve communication, user confidence, and overall training effectiveness.
Research objectives:
Evaluate the functional accuracy of intent recognition systems.
Does it understand and execute user commands and requests?
Asses the usefulness and quality of real-time AI translation features.
Does it contribute positively to team-based interactions among multilingual users?
Examine the role of AI assistance in enhancing task and training performance, as well as user experience.
Do users perceive the AI assistant as helpful and supportive in achieving training objectives?
Verify the inclusivity and accessibility benefits of integrating real-time translation in multi-cultural training contexts.
Does it reduce communication barriers?
How do different language patterns affect the performance?
Analyse user attitudes towards the integration of AI systems (ethics, accuracy, functionality, data protection).
Investigate how the presence of an AI-based Training Partner affects the user autonomy.
Overall assumptions:
02
Solution
How We Measured AI’s Smarts in VR Training
divide and conquer
1
Test: AI-based Training Partner
What we mainly test: Intent Recognition
How: Users participate in a single-person training scenario
2
Test: AI real-time translation
What we mainly test: Translation Efficiency and Impact
How: Users participate in a multi-person training scenario in groups of 2 or 3.
Before I joined, a preliminary survey of over 30 mostly open-ended questions had been created. I began by refining it by analysing each question to ensure clarity, relevance, and measurable value.
My focus was on making questions more direct, reducing bias, and selecting appropriate response types to enable meaningful comparison and reliable analysis of user feedback.
2
language versions - English & German
2
parts: a pre-test questionnaire and a post-test evaluation
aim
to understand how AI intent recognition and translation affect user confidence, communication, inclusivity, and willingness to adopt AI-assisted VR training systems.
pre-test section
gathers demographic information, language background, familiarity with AI and XR technologies, and initial expectations or concerns about AI in training contexts.
post-test section
participants rate aspects such as accuracy, response time, trust, human-likeness, privacy concerns, and perceived usefulness. Open-ended and multiple-choice questions allow participants to share insights about missing features, emotional engagement, and comfort with data processing trade-offs.
While my colleagues recruited participants across five different channels, I focused on preparing clear, user-friendly consent forms and participant handouts with user stories to help them understand the training scenarios they were to complete.
THe Reactor Cleaning VR training
1
Single User Scenario:Users complete a short tutorial on movement and tool use before starting training. As the Cleaner, they follow written and spoken instructions from the AI Assistant to safely shut down, open, and clean the reactor.
2
Multi-User Scenario:
Three roles work together - Operator, Cleaner, and Assistant.The Operator monitors the dashboard, gives verbal instructions, and confirms task progress.The Cleaner performs the reactor cleaning, following the Operator’s directions and communicating through a wrist walkie-talkie.The Assistant supports the Cleaner by handing tools and coordinating with both teammates.

01
Single User story

02
Multi-User sample story for Operator
Can we clean the reactor?
1 Trying out the app
Users participated in 45 minute sessions, completing the tasks in the VR scenario and filling out the surveys.
Live observations and follow-up questions proved useful.
What we prepared:


2 Application Audit
Additionally, I made many observations during the tests and identified points of improvement for designing training applications.
03
Results
Reality Check: When AIXTRA Met Real Users
Methodology:
36 participants
16 users tested the single-user scenario.
24 users tested the multi-user scenario in groups of 2 or 3.Diverse linguistic backgrounds, including native and non-native speakers, different age groups, and various experiences with AI and XR technology.
Procedure
Participants completed structured tasks in the VR scenario.
Quantitative error metrics were collected through Unity logs. Quantitative and qualitative user feedback was collected via a survey to capture real-time performance and user experience.
Data collection
Survey feedback (quantitative & qualitative), direct observation.
Used Likert scales and multiple-choice questions to evaluate.







Key findings:
Intent Recognition:
Overall functional, with correct interpretations in most cases, but less accurate with ambiguous or colloquial phrasing.
Translation Accuracy:
High in simple contexts; reduced accuracy in domain-specific or idiomatic expressions. Occasional mistranslations altered the meaning.
User Experience:
Participants valued speed and accessibility. However, lack of nuance in translations and occasional contextual mismatches caused confusion.
User Concerns :
While users recognise the practical benefits of AI-assisted training systems, a significant portion voice concerns about data privacy, potential algorithmic bias, and the reliability of AI responses.
User study evaluation and improvement points:
Misalignment Between Users & Target Audience
Test participants lacked the domain knowledge expected by the app, leading to confusion.
What I suggest
Ensure test users are either pre-briefed or provided with clear, accessible materials (visual aids, simple instructions, minimal jargon) before using the application.
AI Assistant’s Lack of Contextual Awareness
Even when users sought help, the AI assistant could not assess the user’s current situation or reason about what support was needed.
Assistance was only effective when highly specific questions were asked.
What I suggest
Provide clear definitions, images, or tutorial content for key equipment.
Lack of Clear Feedback & Completion Cues
Users were unsure whether tasks were correctly completed.
They expected the assistant to monitor their progress like a human assistant might, but the assistant only reacts to explicit queries.
The application's interface misleadingly implies that the assistant is actively involved, fostering false expectations about its capabilities.
What I suggest
Set realistic expectations about the assistant’s role and provide more explicit task progression indicators and completion confirmations.
Breakdown in Multi-User Scenarios
Operators unfamiliar with procedures caused confusion, amplified by weak translations.
What I suggest
For multi-user tests, only assign the operator role to someone already familiar with the procedure. Let the cleaner and assistant rely on this knowledgeable operator to simulate the intended dynamic.
Maintaining the Learning Value of the VR Experience
Although users are operating in a safe, simulated environment, the goal is still effective learning.
One user mentioned that due to frequent back-and-forth with the assistant and translation issues, she wasn’t confident she'd retained the full process after completing the training.
Users often didn’t finish tasks or learn anything due to accumulated friction - from confusing tools, unclear instructions, or non-intuitive controls. These obstacles discouraged exploration and led to avoidance of the assistant or key functionality.
My insight
Cognitive overload, caused by translation friction and repeated clarification, can erode learning outcomes -even if the task is eventually completed.
the Role and Perception of the AI Assistant
Users were unaware of the AI assistant’s capabilities, leading to underuse of features like object highlighting. Misaligned expectations also emerged, as many assumed the assistant would act like a human colleague.
My recommendation
Explicitly communicate the assistant's capabilities.
Make key features more obvious (e.g. blinking highlights, greying out irrelevant elements).
Consider offering onboarding prompts and expanding the tutorial demonstrating assistant interactions.
A few notes on translation feature testing:
Evaluating Translation Within the Training Context
Users remained alert to translation errors, questioning odd or unsafe phrases like “breathe inside the reactor.” This critical attitude shows strong engagement and situational awareness — a positive indicator that translation issues did not compromise understanding or safety.
My insight
Users are capable of detecting mistranslations. However, subtle mistranslations may go unnoticed, which could affect task accuracy or safety.
The Role of the Operator in Translation Success
When operators were unsure of the procedure, their instructions became lengthy, hesitant, incomplete or fragmented. This ambiguity led to lower translation quality and confusion downstream.
Principle
Clear source communication = better translation quality. Operators should be well-trained or scripted to ensure consistent, translatable instructions.
Translation Shortening and Information Loss
The translation system often condensed long, rambling operator responses, which improved clarity.
However, there is a risk that key information may be omitted during this shortening process.
My recommendation
Implement post-session checks to evaluate whether essential information is preserved in translated responses.

Users become frustrated when basic interactions don’t match real-world expectations. For example, in the wrench tutorial, they turned the tool incorrectly because the interface didn’t mimic real-world behaviour. Interfaces should align with real-world expectations and muscle memory.
In physical training, a mentor guides learners step by step, but in VR this guidance must be replaced with tools like manuals, diagrams, or visual cues. Current implementations, such as hovering instruction boards, are often cluttered and poorly positioned. Provide just-in-time instructions placed at eye level or within reach to make guidance intuitive and easily accessible.
Process:
Thanks for training with us!Care for projects with more UI focus?

on a mission:
Helping people train, learn, and collaborate across languages — no matter where they are (or which reality they're in).
“A bit of a nightmare,” they say — and I’m instantly intrigued. A casual chat about a user study turns into me joining the project at full speed — right near the finish line, of course. We scrap the old plan, start from zero, and race the clock, and turn chaos into clarity.
Process:
00
Overview
Teaching AI to Train Us
About the application
AIXTRA is a VR training application developed within the EU-funded VOXReality initiative. It integrates real-time AI translation to overcome language barriers and an AI-based virtual training partner to simulate human-to-human collaboration.
Info
Project type: VR Application audit and testing
Medium: VR Training
Role: UX Researcher
Team: 11
Problem
AI systems are increasingly deployed in training environments, but concerns remain about their ability to correctly recognise user intent, provide accurate translations, and build user trust. This study explored these issues through two targeted experiments.
Outcome
User Study:
Total participants: 36
Test sessions: total 24 sessions across 6 days.
01
problem
Does AI Really Get What You Mean — and Help You Learn?
When I joined theproject, the goals were undefined and the requirements unclear. The VR training app has never been tested by anyone other than the developers, and it was built intuitively, without a foundation in UX or UI principles. It’s a tough nut to crack — but also a great opportunity to apply my skills and, in the long run, help the developers shift their perspective for future projects.
It quickly became clear that before testing could begin, we needed to take a step back — define who the users really were, and what problems VOXReality's AI technology was meant to solve.
The app’s core functionalities:
AI-Based Virtual Training Partner
Real-Time AI Translation
The project’s limitations:
1
The app has been already developed without clear user test requirements in mind.
2
No possibilities to test with a control group.
3
Arbitrary requirement to perform a quantitative study with 30 participants.
4
Limited time & resources.
Opportunity:
The AIXTRA app was build to provide a practical use case for presenting improved interaction between humans and AI in XR environments. AIXTRA aims to enhance VR training by overcoming language barriers and providing real-time guidance through a pretrained VOXReality AI-driven virtual assistant and translation system.
GoalS:
The tests aim to determine whether intent recognition, live translation, and AI support can improve communication, user confidence, and overall training effectiveness.
Research objectives:
Evaluate the functional accuracy of intent recognition systems.
Does it understand and execute user commands and requests?
Asses the usefulness and quality of real-time AI translation features.
Does it contribute positively to team-based interactions among multilingual users?
Examine the role of AI assistance in enhancing task and training performance, as well as user experience.
Do users perceive the AI assistant as helpful and supportive in achieving training objectives?
Verify the inclusivity and accessibility benefits of integrating real-time translation in multi-cultural training contexts.
Does it reduce communication barriers?
How do different language patterns affect the performance?
Analyse user attitudes towards the integration of AI systems (ethics, accuracy, functionality, data protection).
Investigate how the presence of an AI-based Training Partner affects the user autonomy.
Overall assumptions:
02
Solution
How We Measured AI’s Smarts in VR Training
divide and conquer
1
Test: AI-based Training Partner
What we mainly test: Intent Recognition
How: Users participate in a single-person training scenario
2
Test: AI real-time translation
What we mainly test: Translation Efficiency and Impact
How: Users participate in a multi-person training scenario in groups of 2 or 3.
Before I joined, a preliminary survey of over 30 mostly open-ended questions had been created. I began by refining it by analysing each question to ensure clarity, relevance, and measurable value.
My focus was on making questions more direct, reducing bias, and selecting appropriate response types to enable meaningful comparison and reliable analysis of user feedback.
2
language versions - English & German
2
parts: a pre-test questionnaire and a post-test evaluation
aim
to understand how AI intent recognition and translation affect user confidence, communication, inclusivity, and willingness to adopt AI-assisted VR training systems.
pre-test section
gathers demographic information, language background, familiarity with AI and XR technologies, and initial expectations or concerns about AI in training contexts.
post-test section
participants rate aspects such as accuracy, response time, trust, human-likeness, privacy concerns, and perceived usefulness. Open-ended and multiple-choice questions allow participants to share insights about missing features, emotional engagement, and comfort with data processing trade-offs.
While my colleagues recruited participants across five different channels, I focused on preparing clear, user-friendly consent forms and participant handouts with user stories to help them understand the training scenarios they were to complete.
THe Reactor Cleaning VR training
1
Single User Scenario:Users complete a short tutorial on movement and tool use before starting training. As the Cleaner, they follow written and spoken instructions from the AI Assistant to safely shut down, open, and clean the reactor.
2
Multi-User Scenario:
Three roles work together - Operator, Cleaner, and Assistant.The Operator monitors the dashboard, gives verbal instructions, and confirms task progress.The Cleaner performs the reactor cleaning, following the Operator’s directions and communicating through a wrist walkie-talkie.The Assistant supports the Cleaner by handing tools and coordinating with both teammates.

01
Single User story

02
Multi-User sample story for Operator
Can we clean the reactor?
1 Trying out the app
Users participated in 45 minute sessions, completing the tasks in the VR scenario and filling out the surveys.
Live observations and follow-up questions proved useful.
What we prepared:


2 Application Audit
Additionally, I made many observations during the tests and identified points of improvement for designing training applications.
03
Results
Reality Check: When AIXTRA Met Real Users
Methodology:
36 participants
16 users tested the single-user scenario.
24 users tested the multi-user scenario in groups of 2 or 3.Diverse linguistic backgrounds, including native and non-native speakers, different age groups, and various experiences with AI and XR technology.
Procedure
Participants completed structured tasks in the VR scenario.
Quantitative error metrics were collected through Unity logs. Quantitative and qualitative user feedback was collected via a survey to capture real-time performance and user experience.
Data collection
Survey feedback (quantitative & qualitative), direct observation.
Used Likert scales and multiple-choice questions to evaluate.







Key findings:
Intent Recognition:
Overall functional, with correct interpretations in most cases, but less accurate with ambiguous or colloquial phrasing.
Translation Accuracy:
High in simple contexts; reduced accuracy in domain-specific or idiomatic expressions. Occasional mistranslations altered the meaning.
User Experience:
Participants valued speed and accessibility. However, lack of nuance in translations and occasional contextual mismatches caused confusion.
User Concerns :
While users recognise the practical benefits of AI-assisted training systems, a significant portion voice concerns about data privacy, potential algorithmic bias, and the reliability of AI responses.
User study evaluation and improvement points:
Misalignment Between Users & Target Audience
Test participants lacked the domain knowledge expected by the app, leading to confusion.
What I suggest
Ensure test users are either pre-briefed or provided with clear, accessible materials (visual aids, simple instructions, minimal jargon) before using the application.
AI Assistant’s Lack of Contextual Awareness
Even when users sought help, the AI assistant could not assess the user’s current situation or reason about what support was needed.
Assistance was only effective when highly specific questions were asked.
What I suggest
Provide clear definitions, images, or tutorial content for key equipment.
Lack of Clear Feedback & Completion Cues
Users were unsure whether tasks were correctly completed.
They expected the assistant to monitor their progress like a human assistant might, but the assistant only reacts to explicit queries.
The application's interface misleadingly implies that the assistant is actively involved, fostering false expectations about its capabilities.
What I suggest
Set realistic expectations about the assistant’s role and provide more explicit task progression indicators and completion confirmations.
Breakdown in Multi-User Scenarios
Operators unfamiliar with procedures caused confusion, amplified by weak translations.
What I suggest
For multi-user tests, only assign the operator role to someone already familiar with the procedure. Let the cleaner and assistant rely on this knowledgeable operator to simulate the intended dynamic.
Maintaining the Learning Value of the VR Experience
Although users are operating in a safe, simulated environment, the goal is still effective learning.
One user mentioned that due to frequent back-and-forth with the assistant and translation issues, she wasn’t confident she'd retained the full process after completing the training.
Users often didn’t finish tasks or learn anything due to accumulated friction - from confusing tools, unclear instructions, or non-intuitive controls. These obstacles discouraged exploration and led to avoidance of the assistant or key functionality.
My insight
Cognitive overload, caused by translation friction and repeated clarification, can erode learning outcomes -even if the task is eventually completed.
the Role and Perception of the AI Assistant
Users were unaware of the AI assistant’s capabilities, leading to underuse of features like object highlighting. Misaligned expectations also emerged, as many assumed the assistant would act like a human colleague.
My recommendation
Explicitly communicate the assistant's capabilities.
Make key features more obvious (e.g. blinking highlights, greying out irrelevant elements).
Consider offering onboarding prompts and expanding the tutorial demonstrating assistant interactions.
A few notes on translation feature testing:
Evaluating Translation Within the Training Context
Users remained alert to translation errors, questioning odd or unsafe phrases like “breathe inside the reactor.” This critical attitude shows strong engagement and situational awareness — a positive indicator that translation issues did not compromise understanding or safety.
My insight
Users are capable of detecting mistranslations. However, subtle mistranslations may go unnoticed, which could affect task accuracy or safety.
The Role of the Operator in Translation Success
When operators were unsure of the procedure, their instructions became lengthy, hesitant, incomplete or fragmented. This ambiguity led to lower translation quality and confusion downstream.
Principle
Clear source communication = better translation quality. Operators should be well-trained or scripted to ensure consistent, translatable instructions.
Translation Shortening and Information Loss
The translation system often condensed long, rambling operator responses, which improved clarity.
However, there is a risk that key information may be omitted during this shortening process.
My recommendation
Implement post-session checks to evaluate whether essential information is preserved in translated responses.

Users become frustrated when basic interactions don’t match real-world expectations. For example, in the wrench tutorial, they turned the tool incorrectly because the interface didn’t mimic real-world behaviour. Interfaces should align with real-world expectations and muscle memory.
In physical training, a mentor guides learners step by step, but in VR this guidance must be replaced with tools like manuals, diagrams, or visual cues. Current implementations, such as hovering instruction boards, are often cluttered and poorly positioned. Provide just-in-time instructions placed at eye level or within reach to make guidance intuitive and easily accessible.
Thanks for training with us!Care for projects with more UI focus?