Session Information
10 SES 07 C, Teacher Agency in the Age of AI: Negotiating Pedagogy, Knowledge, and Professional Practice
Paper Session
Contribution
The incorporation of generative AI (GenAI) into physics education embodies a duality of extraordinary potential and considerable educational risk. Recent research indicates that models such as GPT-4o may surpass the performance of typical undergraduates on standardized physics concept inventories; however, this superiority fosters a misleading appearance of mastery (Kortemeyer et al., 2025). A thorough review of the literature uncovers significant concerns related to multimodal processing, the validity of assessments in laboratory environments, and the dependence on students' prior knowledge to alleviate AI hallucinations (Kortemeyer et al., 2025; Nikolic et al., 2025).
First, AI doesn't work the same way in all areas and modes. LLMs are excellent at conceptual tasks that involve text, and they often do better than humans at things like thermodynamics and relativity. However, they are not very good at interpreting visual data like graphs and diagrams. Also, there is a big "language gap" because AI models don't work the same way in all languages. They work better in English and Western European languages but worse in others. This could make global educational inequalities worse (Kortemeyer et al., 2025).
Second, the rise of GenAI puts the accuracy of traditional lab tests at risk. The standard lab report, which tests cognitive skills like writing and data synthesis a lot, is very easy for AI to copy. Conversely, psychomotor objectives (equipment handling) and affective objectives (teamwork, ethics) are significantly more impervious to AI simulation; however, educators presently exhibit diminished confidence in evaluating these areas. To preserve assessment validity, it is imperative to transition from unsupervised written reports to a variety of assessment modalities, such as practical examinations, interviews, and direct observations, which are not easily replicable by AI (Nikolic et al., 2025).
Lastly, the fact that AI can be a "lab partner" shows a paradox of expertise. Case studies with high school students show that in order to use AI to solve problems well, students need to already know enough about the subject to be able to spot "hallucinations" or wrong physics explanations given by the model. Students who are new to a subject and don't have this basic knowledge are more likely to believe analogies that sound good but are scientifically wrong. This can make misunderstandings worse instead of fixing them. So, to use GenAI in classes, we need more than just access to the technology (Kilde-Westberg et al., 2025).
In this context, the emphasis on science teacher training is critical, particularly in relation to GenAI. Programs should not only use GenAI for productivity but also encourage critical evaluation and the development of irreplaceable skills. This design-oriented qualitative study was conducted in a new undergraduate course, Experiments in Primary Science Education. The course treated GenAI as a critical tool for hands-on experiments rather than a credible source of information. The hands-on experiments focus on recognizing AI faults, reassessing valid evaluations, and rethinking instructional identity regarding AI. GenAI is not considered reliable. The study shows how GenAI changes notions of reliable knowledge and teaching obligations by redefining knowledge as an epistemic practice impacted by judgment and uncertainty. Overall, it contributes to discussions on how educational research can adapt to AI-enhanced classroom difficulties and potential. Within this framework, the study is guided by the following research questions:
1. How do pre-service teachers evaluate the scientific reliability of GenAI-generated explanations in hands-on science experiments?
2. How does engagement with GenAI influence pre-service teachers’ conceptions of assessment validity in experimental science education?
3. How do these experiences reshape pre-service teachers’ pedagogical positioning and sense of agency in AI-mediated learning environments?
Method
The study adopts a design-based qualitative research approach, enabling iterative alignment between pedagogical design, data collection, and theoretical interpretation. The research was carried out during a six-week instructional module integrated into the “Experiments in Primary Science Education” course at a faculty of education. The participants were pre-service primary and elementary science teachers enrolled in the course. The educational design has three phases that are all connected. In the first phase, participants are expected to perform science experiments that are in line with primary-level curricula, such as floating and sinking. After doing certain experiments, students ask a GenAI system to explain what they saw. These AI-generated answers are deliberately chosen or directed to incorporate plausible yet scientifically erroneous reasoning. Participants are instructed to evaluate AI outputs against their empirical observations and to substantiate their assessments concerning scientific reliability and accuracy. During the second phase, participants are expected to work on redesigning assessments. For the same experimental tasks, they look at several ways to test, such as traditional written lab reports, oral exams, practical demonstrations, and real-time observational rubrics. Students critically evaluated which learning outcomes could be properly assessed in the context of GenAI and which necessitated human judgment, physical involvement, or ethical thinking. The third phase is all about figuring out how they are as a teacher and where they fit in. At the beginning and end of the module, participants are expected to write down their thoughts on the function of GenAI in science teaching and what teachers should do in AI-mediated learning settings. The sources of data are: Reports on comparative assessments in writing of explanations based on AI and experiments Reflective journals and structured written responses Pre- and post-module pedagogical positioning texts Interviews (individual and group) Qualitative content analysis and an inductive-deductive coding technique are planned to be used. Analytical categories concentrated on epistemic evaluation standards, the rationale for assessment validity, and manifestations of teacher agency and accountability. Triangulation of data sources and repeated peer debriefing will be generated to show the reliability of the analysis.
Expected Outcomes
This research aims to enhance physics and science education research by providing an empirically based and critically focused example for the incorporation of GenAI into experimental learning contexts. Instead of only looking at how well or quickly GenAI works, the study hopes to show how using AI changes teachers’ decisions about what to teach, how they evaluate students’ performances, and how they conduct themselves as teachers through the educational environment changes. First, the study is going to clarify how structured, experiment-focused interactions with GenAI facilitate the cultivation of epistemic awareness among pre-service teachers. It is expected that participants would exceed artificial recognition of fluent AI-generated explanations and progressively depend on empirical evidence, causal reasoning, and conceptual coherence in the assessment of scientific arguments. This result would highlight that science education should see GenAI as a tool for critical analysis rather than a definitive source of knowledge. Second, the study seeks to improve GenAI-era assessment validity discussions. The research engages pre-service teachers in a critical review of evaluation methodologies to emphasize a change toward learning outcomes that focus on human qualities like experiential judgment, psychomotor skills, ethical reasoning, and collaborative practices. These findings might help reconfigure AI-mediated assessment methods for efficacy. Third, the study aims to demonstrate shifts in pedagogical positioning, with pre-service teachers increasingly viewing their role as epistemic mediators and responsible decision-makers rather than mere distributors of content. This anticipated modification would impact teacher education programs aiming to prepare future teachers for classrooms filled with AI. The study aims to contribute to educational research by demonstrating how experimental science courses can function as critical platforms for examining the risks and opportunities linked to GenAI. The study underscores epistemic integrity, evaluative accountability, and teacher agency, offering a flexible structure for the incorporation of human-centered AI in scientific education.
References
Kortemeyer, G., Babayeva, M., Polverini, G., Widenhorn, R., & Gregorcic, B. (2025). Multilingual performance of a multimodal artificial intelligence system on multisubject physics concept inventories. Physical Review Physics Education Research, 21(2), 020101. https://doi.org/10.1103/98hg-rkrf Nikolic, S.,Suesse, T.F., Grundy, S., Haque, R., Lyden, S., Lal, S., Hassan, G.M., Daniel S. & Belkina, M. (2025). Assessment integrity and validity in the teaching laboratory: adapting to GenAI by developing an understanding of the verifiable learning objectives behind laboratory assessment selection, European Journal of Engineering Education, 50(4), 673-701, https://doi.org/10.1080/03043797.2025.2456944 Kilde-Westberg, S., Johansson, A., & Enger, J. (2025). Generative AI as a lab partner: A case study. Physical Review Physics Education Research, 21(2), 020119. https://doi.org/10.1103/ggy1-3kjk
Update Modus of this Database
The current conference programme can be browsed in the conference management system (conftool) and, closer to the conference, in the conference app.
This database will be updated with the conference data after ECER.
Search the ECER Programme
- Search for keywords and phrases in "Text Search"
- Restrict in which part of the abstracts to search in "Where to search"
- Search for authors and in the respective field.
- For planning your conference attendance, please use the conference app, which will be issued some weeks before the conference and the conference agenda provided in conftool.
- If you are a session chair, best look up your chairing duties in the conference system (Conftool) or the app.