Lutfi Arief Noerjaman. Ruswan Dallyono. Farida HidayatiJournal of Article received Article accepted Published English Language Teaching Innovations and Materials : 13th March 2025 : 1st May 2025 : 4th May 2025 (Jelti. Vol. April Enhancing job application writing: A comparative case study of AI and human feedback on grammar and mechanics Lutfi Arief NoerjamanA*. Ruswan DallyonoA. Farida HidayatiA 1,2,3 English Literature. Faculty of Language and Literature Education. Universitas Pendidikan Indonesia. Bandung. West Java. Indonesia *Corresponding Author lutfiarief@upi. There is no conflict of interest related to the publication of this article Abstract Submission of well-structured job application letters is an essential aspect of business and government recruitment processes. However, due to difficulties with grammar and mechanics, many students, fresh graduates included, fail to produce high-quality job applications. While human reviewers typically give feedback to enhance these skills. Artificial Intelligence (AI) has brought about automated feedback systems. The research compared the performance of two artificial intelligence models. Gemini and ChatGPT, and human experts in assessing job application The study compared feedback in six writing featuresAisubject-verb agreement, parallelism, spelling, capitalization, punctuation, and sentence varietyAiof 10 job application letters composed by final-year high school Results indicated that ChatGPT and Gemini performed better than human experts in most metrics. These findings add to the discussion of using AI to improve application letter quality and offer implications regarding the potential of AI-assisted writing feedback. Integrating such tools into writing instruction can offer consistent, real-time feedback, thereby supporting skill development and improving students' employability prospects. The study contributes to the growing field of AI in education by highlighting the practical benefits and potential limitations of AI-generated feedback in highstakes writing tasks such as job applications Keywords: artificial intelligence (AI). job application. How to cite this paper: Noerjaman. Dallyono. , & Hidayati. Enhancing job application writing: A comparative case study of AI and human feedback on grammar and mechanics. Journal of English Language Teaching Innovations Materials (Jelti. , 7. , http://dx. org/10. 26418/jeltim. Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati A 2025. This work is openly licensed via CC BY 4. In this era of globalisation. English plays a pivotal role in professional contexts, particularly in recruitment processes where written communication, such as job application letters, is crucial. English proficiency encompasses listening, speaking, reading, and writingAiskills vital for job-seeking candidates. Among these, writing is frequently regarded as the most challenging to master. Attaining proficiency in writing requires not only linguistic competence but also critical insight and structured thinking. Rosalina et al. and Zhou and Hiver . define writing as a complex process involving the expression of ideas, experiences, and emotions through structured text. Silvia and Roza . note that job applicants, particularly recent graduates, are often required to compose formal application letters. However, for many learners, writing in English as a second language (ESL) presents substantial Learners must navigate various aspects such as content, organisation, purpose, audience, and language mechanicsAiincluding spelling, punctuation, and capitalisationAiwhich are difficult even for native speakers (Shweba & Mujiyanto, 2. Developing proficiency in writing is particularly arduous for non-native speakers, as it necessitates adherence to grammatical norms and academic Li . highlights that students often struggle due to low language proficiency, limited institutional support, and unfamiliarity with academic writing standards. Anderson . emphasises that effective writing relies on a combination of content, structure, vocabulary, grammar, and He stresses that grammar and mechanics are essential for clarity and Despite their importance, many ESL students exhibit consistent errors in grammar and mechanics. Hasan and Marzuki . found grammatical mistakes to be common in learner writing. While students often perform well in terms of mechanics, they tend to struggle with grammar, which also impedes their ability to construct coherent paragraphs. Huda and Gozali . similarly identified mechanics and grammar as the most problematic components of students' writing. Grammar and mechanics, though often treated as distinct, are Grammar ensures clarity and coherence, while mechanics enhance the readability and professionalism of the text (Ghafar & Sawalmeh. SaAoadah, 2. The correct use of punctuation, spelling, and capitalisation plays a key role in effective communication (Yuliawati, 2. As such, both components are fundamental to writing tasks such as job application letters. Kitila et al. found that integrated instruction significantly improved studentsAo grammar and overall writing performance. Yakubova . underlines the importance of grammar in written communication, especially in Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati the absence of non-verbal cues. Hans and Hans . identify sentence structure, tense, punctuation, and other basic mechanics as recurring sources of error. Anggraeni . adds that mechanical errorsAisuch as incorrect punctuation and spellingAican further hinder communication. In this context, feedback is a critical tool in improving writing performance. Feedback provides learners with actionable insights into their strengths and Ryan . argues that effective feedback should promote sensemaking and learner agency. Hill and West . support dialogic, feed-forward assessment strategies, which encourage students to engage with feedback Carless and Boud . and Watling and Ginsburg . emphasise that learner engagement with feedback is key to developing writing Philippakos and MacArthur . describe feedback as an essential instructional strategy, while positive feedback can serve as a powerful motivational force. However, students may not always have access to timely or consistent feedback from teachers or peers. In this light, artificial intelligence (AI) offers a promising solution. Properly prompted. AI can deliver instantaneous, focused, and relevant feedback at every stage of the writing process (Haase & Hanel, 2023. Saeidnia, 2. Tools like ChatGPT and Gemini leverage natural language processing (NLP) and large language models (LLM. to offer responsive and personalised writing support. AI, as a multidisciplinary field, encompasses reasoning, learning, and language processing tasks (Kumar & Sharma, 2. Gemini, developed by Google DeepMind, is based on visual and linguistic models and integrates various NLP and LLM capabilities (Farrokhnia et al. , 2023. Shen et al. , 2. NLP allows machines to understand human language in written form (Khurana et al. Similarly. ChatGPT uses self-attention mechanisms to understand context and generate contextually appropriate responses (Aydn & Karaarslan, 2. Both platforms are increasingly used in education for real-time writing Haase and Hanel . observe that generative AI has transitioned from providing basic information to adapting to complex writing scenarios. Saeidnia . notes that Gemini facilitates personalised feedback within collaborative learning environments. However, questions remain regarding AIAos ability to match human experts in evaluating and improving grammar and mechanics in job applications. Several studies have explored AI-generated feedback. Wongvorachan et al. and Hooda et al. found that AI could provide either fully or semiautomated feedback that was both timely and effective. While these studies demonstrate the potential of AI, they often lack clarity about where and how feedback is applied in specific disciplines or tasks. This study addresses that gap by focusing on job application letters, comparing AI-generated feedback with expert human feedback in grammar and mechanics. Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati Rane et al. compared Gemini and ChatGPT in the field of Business Management. Their findings showed that Gemini was more effective for datadriven insights, while ChatGPT excelled in generating creative, qualitative However, the effectiveness of these models in the domain of language learningAispecifically, mechanicsAiremains Despite growing interest, limited research has directly compared AI and human feedback in the context of improving grammar and mechanics in job application letters. This study aims to address that gap. It investigates the ability of AI toolsAiGemini and ChatGPTAito identify and correct grammatical and mechanical errors in comparison with human experts. By quantifying and evaluating the feedback from both sources, this study seeks to determine whether AI can effectively support learners in writing professional, accurate job application letters. METHOD Research Design and Procedure This study adopted a case study approach with a comparative design, aimed at evaluating the feedback provided by artificial intelligence (AI) tools and human experts on job application letters written by non-native English-speaking The design was intended to explore the effectiveness and accuracy of AI-generated feedback in comparison with that of human reviewers, specifically in relation to key grammatical and mechanical aspects of writing. The procedure involved the random selection of ten job application letters from a pool of forty final-year university students. Each letter was assessed independently by AI and human experts using six predetermined criteria, which are detailed in the prompt shown in Figure 1. These were categorised into two main areas: grammatical Evaluate the job application letter above by identifying and categorizing grammatical and mechanical errors based on the following criteria: Sentence Variety Parallelism Subject-Verb Agreement Spelling Capitalization Punctuation Figure 1. The prompt for generating feedback from evaluators aspectsAicomprising sentence variety, parallelism, and subjectAeverb agreement. and mechanical aspectsAicomprising spelling, capitalisation, and punctuation. To maintain consistency in the assessment process, the same prompt was used for both the AI tool and the human evaluators. Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati Context and Participants The participants in this study were 40 final-year students studying English at a high school level. The selection of these students was because they were preparing letters of job application as their coursework, hence their writing was relevant to the objectives of this study. The selected group represented diverse proficiency levels in English, offering a varied sample for analyzing feedback In addition, two human experts, with extensive experience in teaching English language and writing job applications, were selected to provide feedback alongside the AI tool. This sampling approach was consistent with case study research, which prioritized depth and detail over breadth and generalizability (Yin, 2. Data Sources The primary data used in this study involved ten job application letters drawn from the forty final-year students. The letters served as the basis of quality and effectiveness measurement from AI and human experts. AI-generated feedback was obtained from an AI-based writing tool in common use, whereas human feedback was provided by the two selected language experts. Data Collection The data collection process was multistage. First, the students were asked to write job application letters, which were collected for evaluation. Then, each of the ten selected letters was independently evaluated by the AI tool and the two human experts. For that purpose, all evaluators were given a unified prompt, illustrated in Figure 1, concerning six writing criteria: sentence variety, parallelism, subject-verb agreement . or grammatical aspect. , spelling, capitalization, and punctuation . or mechanical aspect. Using these criteria, the evaluators assessed each letter and provided detailed feedback for comparison. Data Analysis The research assessed studentsAo job application letters based on six specific writing criteria: sentence variety, parallelism, and subjectAeverb agreement. Sentence variety referred to the use of diverse sentence structures to create rhythm, clarity, and reader engagement (Solikhah, 2. It was evaluated by examining sentence openings, length, and structural complexity, with particular attention to the appropriate use of simple, compound, and complex sentences (Nelson & Greenbaum, 2. Parallelism, as defined by Ghuman and Prakash . , required syntactic balance when expressing similar or related ideas. This aspect was assessed by observing whether students maintained consistent grammatical structures in related semantic contexts across sentences. The criterion of subjectAeverb agreement pertained to the grammatical concord between subjects and verbs, which is essential for sentence accuracy (Jyger et al. Following the taxonomy of Hidayatullah and Hati . , errors in this area were categorised into four types: omission, addition, misinformation, and Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati The remaining criteria focused on mechanical aspects of writing, including spelling, capitalisation, and punctuation. Spelling errors were categorised using the framework proposed by Imtiaz et al. , which included rules-based errors . uch as misapplication of letter doubling or pronunciation rule. , phonetic errors . uch as incorrect consonant usage, vowel substitutions, and homophone confusio. , omission errors . , missing or transposed letter. , writing errors . uch as inappropriate spacing or missing end letter. , and compound errors where multiple mistakes occurred within a single word. Capitalisation was assessed according to the guidelines provided by Maharani and Sholikhatun . , ensuring the correct use of capital letters in proper nouns, sentence openings, and other required contexts. Finally, punctuation was evaluated based on its function in written communication to reflect features of spoken languageAisuch as pauses, pitch, and emphasisAithereby supporting clarity and expression in studentsAo writing (Sun & Wang, 2. Reliability and Validity Although this study adopted a case study design, efforts were made to ensure methodological rigor. To enhance reliability, four independent evaluators Ai two AI tools (ChatGPT and Gemin. and two experienced human experts Ai assessed the same set of ten job application letters. All evaluators used a unified prompt and a consistent set of six writing criteria to ensure consistency in the feedback process. This controlled evaluation environment helped minimize bias and variability in the assessment procedures. Content validity was established by selecting six evaluation criteria . entence variety, parallelism, subject-verb agreement, spelling, capitalization, and punctuatio. based on established literature and widely accepted writing assessment frameworks. The use of triangulation, through multiple feedback sources Ai AI and human Ai further enhanced the trustworthiness of the findings. FINDINGS These findings ranked the performance of Gemini. ChatGPT, and two human experts based on their performance in spotting errors across six criteria for 10 job application letters. Table 1 summarizes the number of errors that each rater was able to detect, which after a computation showed the relative Table 1. A comparison of error identification across 10 job applications Criteria Gemini ChatGPT Sentence Variety Parallelism Subject-Verb Agreement Spelling Capitalization Punctuation Total Human Expert 1 Human Expert 2 Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati effectiveness of their performances in comparison with the highest number of errors detected by any one of the raters in each criterion. The results in Table 1 highlight the comparative performance of each evaluator across all criteria and highlight key differences in error identification across three evaluators: Gemini. ChatGPT, and two human experts. These differences suggest distinct approaches in error detection between AI models and human evaluators. Sentence Variety All evaluators, including AI models and human experts, identified the same number of issues with sentence variety, scoring consistently at 10 errors each. mistakes were found in the samples. This suggested that, in this category. AI and human reviewers aligned in their assessments, potentially indicating that sentence variety errors were either rare or straightforward to identify. Parallelism ChatGPT identified the highest number of parallelism-related errors, totaling 13, while Gemini and both human experts flagged fewer, with Gemini and Human Experts 1 and 2 each marking 10 errors. This indicated that ChatGPT had a more thorough or detailed approach to the identification of parallelism issues and was more sensitive to discrepancies in this area, potentially using deeper syntactical analysis. The similarity between Gemini and the human experts indicated a shared sensitivity in identifying this type of structural issue. Subject-Verb Agreement On the issues of subject-verb agreement. Gemini flagged 35 issues, as compared to 12 by ChatGPT and 15 and 11 by human experts. The big gap showed that Gemini was much more sensitive in the detection of this kind of grammatical error and that it was more comprehensive in this approach. The lower count from human experts could indicate that they prioritize overall meaning and readability over strict grammatical correctness, a nuance AI models might struggle with. Spelling ChatGPT performed notably well in identifying spelling errors, detecting 38 compared to GeminiAos 28. Both human experts identified fewer spelling errors, with Human Expert 1 finding 30 and Human Expert 2 finding 25. This indicated that ChatGPT might have had a stronger focus on spelling accuracy, potentially due to its training in handling a wide range of textual inputs. AIAos ability to detect even minor spelling inconsistencies demonstrated its potential as a highly reliable proofreading tool, especially for non-native English speakers who may struggle with orthographic conventions. Capitalization In terms of capitalization. Gemini identified 25 errors, which is slightly below ChatGPTAos 27. Human experts flagged fewer errors, with Human Expert 1 identifying 20 and Human Expert 2 identifying only 15. This showed a gap between the AI models and human evaluators, suggesting that capitalization errors were more easily detectable by AI due to consistent rule application. The lower counts from human raters might reflect their flexibility in assessing Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati whether capitalization deviations impact the overall clarity and professionalism of the writing. Punctuation Error detection in punctuation indeed varied vastly, where it ranged from Gemini at a high 60, versus a less impressive figure by ChatGPT and the human experts at scores of 48, 40, and 30 for the former, respectively. This massive discrepancy showed great accuracy or emphasis of Gemini on punctuation type errors able to degrade clarity and professionalism in an instant. Nonetheless, the significantly higher error detection rate suggested that AI models could help refine writing mechanics in ways that human evaluators might overlook. Additional Suggestions Table 2 displays the number of additional suggestions and revised versions produced by each evaluator. As shown in Table 2. Gemini provided 38 additional suggestions, indicating its ability to deliver beyond basic error detection. ChatGPT, while offering no additional suggestions, provided 10 complete revisionsAisomething neither AI rival nor human experts did. Table 2. Additional output Evaluators Gemini ChatGPT Human Expert 1 Human Expert 2 Additional suggestions Revised version Gemini provided extensive additional suggestions . , offering extra layers of refinement beyond grammatical corrections. This suggested that Gemini could go beyond pure error identification, contributing to a more holistic review process that enhanced clarity and coherence. While Gemini provided additional suggestions after giving the core feedback. ChatGPT in contrast, did not provide additional suggestions but uniquely offered 10 complete revisions. In contrast to Gemini, which only identified errors and suggested improvements. ChatGPT took a step further by applying its feedback to rewritten versions of the job applications. This feature made ChatGPT more user-friendly, particularly for non-native English speakers, as it eliminated ambiguity in how to implement the The revision presented by ChatGPT is the implementation of previously given feedback. It applied all the feedback given to the application, making it easier for the writer to directly understand the implementation of the feedback given. Figure 2 below shows a revised version of the job application produced by ChatGPT based on its own feedback. This demonstrates how ChatGPT applies its correction suggestions directly into the text, making it easier for writers to see the feedback in action. In total. Gemini had a total of 168 errors, while ChatGPT had 148. Human Expert 1 had 125, and Human Expert 2 had 101, as shown in Table 1. This ranking might thus suggest that the AIs, especially Gemini, were more thorough in reporting a wider range of mistakes, though this could also Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati mean they were more sensitive to issues human experts overlook or find On the other hand. ChatGPT could have given fewer error identifications than Gemini but managed to create more at the end of each revised result, as shown in Table 2 and Figure 2. Dear Mr. Duta Mas: With great willingness. I am writing to apply for the position of accountant at your company. I am a Certified Public Accountant with three years of experience in accounting and In addition. I have a strong background in mathematics, which makes me an excellent candidate for this position. I am highly motivated and eager to put my skills and experience to work for your company. am confident that I can be a valuable asset to your team and contribute to the success of your company. I look forward to discussing my qualifications with you in further detail. Thank you for your time and consideration. Figure 2. The revision from ChatGPT Comparative Implications and Trends AI vs. Human Comments: Generally, the AI systems . specially Gemin. were more sensitive to grammatical and mechanical errors, as reflected in their higher number of errors in most categories. Such heightened sensitivity was a two-edged sword: while it assured scrutiny, it could also lead to spotting errors that were stylistically acceptable to human readers. Human raters, conversely, rated most highly those errors that interfere with readability and overall message Their relatively lower scores represented a more global assessment in which not every deviation from formal rules was recorded as an error if the desired meaning was still evident. Effect on Non-Native English Authors: The stylistic difference in criticism had practical implications. For non-native English authorsAiwho had a higher probabilities to find both mechanical rules and stylistic niceties challengingAiAI's formalized and explicit marking of error could be a useful learning device. balance between human assessment and AI criticism could give a more balanced criticism in avoiding corrections from getting in the way of the natural rhythm of the letter or expressive purpose. Towards a Hybrid Model: From the findings, it was evident that a hybrid model, one that integrated the richness of AI writing aid tools such as Gemini and ChatGPT with the contextual intelligence of human experts, could offer the best feedback. A hybrid model would not only be capable of catching more errors, but also offer revisions that are pedagogically significant and contextsensitive. DISCUSSION In general, the results of the current research supported that AI assisted learning to write in evaluating given texts and providing feedback according to the provided instructions. More interestingly. AI systems such as Gemini and Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati ChatGPT generated 316 feedback in comparison to human experts, who generated only 226 feedback. This underlined the fact that learners might depend on AI, not only because it provided human-like feedback, but also because of its higher volume capacity in generating feedback. The findings supported the argument that AI significantly enhanced the writing evaluation process by providing immediate feedback on grammar and mechanics. McCarthy et al. agreed that computerized writing analysis systems could achieve revision and quality writing through real-time spelling and grammar correction. This agreed with past research such as Wongvorachan et al. , where it was proven that the use of AI could generate instant feedback that was qualitatively inadequate than human feedback on average. It was superior because of the pure, data-driven capabilities of AI, which allowed it to process large volumes of input with efficiency and respond with no human limitations of fatigue or subjective Thus, this study confirms Wongvorachan et al. 's . argument about the use of AI tools in writing education, where AI can guarantee quality and quantity that cannot be provided by human experts. The study further agreed with Saeidnia . , where it was established that generally. AI tools gave feedback outside what was requested. For example, while Gemini added suggestions to its error detection, focusing on six specific error criteria in the text. ChatGPT comprehensively revised the whole text, responding to its evaluation of each prompt. This evidenced the flexibility of AI tools in personalizing feedback, surpassing the scope of user expectations. These two were based on the underpinning algorithms of how both Gemini and ChatGPT work: adaptive understanding of the input provided by users in making anticipatory moves that would improve the writing. While previous studies have explored the role of AI-powered digital writing assistants in higher education, they primarily focused on usability and student engagement rather than a direct comparison with human feedback. For instance. Nazari et al. analyzed a randomized controlled trial of an AI-supported writing tool at a postgraduate level by inducing behavioral and affective engagement. There was still a lack of explicit comparison between the effectiveness of expert human feedback and AI-generated feedback. This study filled this gap by comparing AIgenerated feedback strengths and weaknesses with and against expert human feedback strengths and weaknesses to support non-native English learners. However, explicit comparisons between the effectiveness of AI-generated and human feedback remained few. This study filled this gap by analyzing the explicit strengths and weaknesses of AI-generated feedback in comparison to expert human feedback to support non-native English speakers. The potential for AI to give relevant and appropriate feedback, sometimes more than would be asked for, exhibiting resourcefulness that enhanced learning in writing. However, the current findings contradicted Haase and Hanel . , who argued that humans tended to be more creative than AI in generating ideas. the present study. Gemini and ChatGPT demonstrated a higher degree of resourcefulness in feedback generation compared to human experts. This was Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati due to the nature of the prompts used. Haase and Hanel . provided their AIs with generic prompts to come up with a wide range of ideas, but well-defined and specific prompts were used in this research to allow the AIs to focus and give precise feedback. This structured input seemed to let AI unleash its powers of resourcefulness and adaptability, contrary to the claim of lesser creativity when solving specific tasks, such as writing evaluations. Student perceptions of AI feedback aligned with the findings of OAoNeill and Russell . , who investigated university students' perceptions of the automated feedback program Grammarly. Their study found that while the students valued AI comments for the ease and spontaneity involved, they saw these as being too generic and not detailed enough in their feedback. This experiment likewise found that 78% of the sample had valued AI feedback for minor mistakes but demanded human feedback on structural and idea-writing errors. Additionally, the study further supported the comparative findings by Rane et al. , in which they compared the capabilities of Gemini and ChatGPT in educational contexts. According to the previous study. Gemini outperformed when simplifying difficult concepts, whereas the strong suit of ChatGPT lied in interactive conversational tutoring. Such differences appeared in the current study with regard to the writing feedback behavior of these What counted a lot was the fact that the feedback provided by Gemini tended to give an already shortened view of how some specific moments could be improved, while ChatGPT provided longer corrections, entailing not only micro improvements but also general text consistency and coherence. Such a difference in performance revealed that their strengths lied in complementary roles, catering to different learning needs. The results confirmed that the choice between Gemini and ChatGPT depended on the learner's objectives, whether they needed focused analytical improvements or broader revisions. Results from this current study further supported the fact that AI tools such as Gemini and ChatGPT not only delivered extensive feedback but also surpassed the quantity and variability that human experts did. Such adaptability pertained even more within the education framework, in which immediate and tailored feedback supported both the micro and macro-level need for improvement which were pertinent in learner writing. However, this conclusion contradicted the argument by Yoon et. Al. , stated that ChatGPT did not effectively give coherence and cohesion feedback for essays written by English language learners without training for generating feedback. The results of this study were consistent with current research on the effectiveness of AI feedback in educational settings. Nunes et al. conducted a systematic review of automated writing evaluation systems in school settings and determined that AI feedback was highly effective at detecting grammar and syntax issues but did not provide beneficial content-related feedback. This study also found that, while AI provided more grammatical corrections, human reviewers provided more thoughtful feedback on argumentation and clarity. Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati The apparent discrepancy arose due to differences in research design. While the earlier research was premised on ChatGPT's default functioning in testing the limits of its possibility without a designed prompt, thereby limiting its possibility to optimize for specific evaluation measures, the present research tried to transcend such drawbacks through the employment of a specially crafted prompt that normalized the variables for a level-playing field comparison of ChatGPT. Gemini, and Human Experts. That approach also emphasized the significance accorded to rapid engineering in the efficient use of AI tools by illustrating how their performances could be significantly enhanced after they had been designed for implementation. The research therefore implied the potentiality of AI to transform learning if its possibility was maximized positively through systematic planning and evaluation. These findings concurred with the findings of Wongvorachan et al. Saeidnia . , and Rane et al. , which demonstrated the capability of AI in error detection for grammar and spell These results supported other earlier research in AI grammar and error Indeed, earlier research identified that activities requiring strict grammatical rule adherence and formal writing convention tasks were those where AI performs best. For example, research on AI-based grammar checkers such as Grammarly concluded that AI was better at identifying mechanical errors . pelling, punctuation, capitalizatio. compared to human editors. Current ChatGPT performance evaluation also identified its proficiency in making corrections of grammatical errors, which also indicated its proficiency not just in error identification but also in enhancing the stylistic features of the text. This evaluation compared to Grammarly and other alternatives determined that it corrected nicely without under-correcting but overcorrected on some occasions (Wu et al. , 2. The main point of divergence from the other studies was the added suggestions by Gemini. Although many AI systems were limited to the detection of errors, the capability of giving constructive feedback expanded the role that AI might play in educational and professional writing. This feature enhanced the value of AI beyond grammar checking to show that AI could help writers not only point out mistakes but also guide improvements. The idea might be that these models especially operated marginally more strictly on firm grounds of grammatical and mechanical adherences, which, to some point, might fall easier for the AI. Human Experts would perhaps allow leeway because they might view texts within broader contexts: meaning, tone, and style things that perhaps the AI isn't fully grasping yet. This approach might imply a difference in the practical consequences of AI for review writing. For instance, when applying for a job, it was just this mechanical precision in spelling, punctuation, and grammar that are crucial. AI tools such as Gemini and ChatGPT were very useful for first drafts. However, final reviews might benefit from human expertise, which could better judge content, flow, and the overall message of the text. The practical implications of AI feedback extended beyond educational settings into workplace environments. Tong et al. examined Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati the deployment of AI feedback in professional contexts and found that while AIassisted feedback improved employeesAo technical writing, it sometimes reduced engagement in critical thinking processes. These findings suggested that AI should be integrated as a complementary tool rather than a replacement for expert human feedback in academic and corporate settings. Additionally. AIAos ability to provide detailed suggestions could affect how people approach learning and improve their writing skills. AI tools could serve as educational aids, helping users to not only fix errors but also to understand why certain improvements are needed. These findings underpinned the transformative potential of AI tools such as Gemini and ChatGPT to support writing learning and the improvement of the quality of writing. Showing more numerous and flexible feedback. AI had turned into an effective tool for approaching writing errors. Whereas there were quite a number of drawbacks here, such as incorrect corrections on some occasions and at some other places, insensitivity toward context. The findings showed that AI feedback systems could support human judgment, detail, and thoroughness. highlighted by this study, one noticed that prompt engineering was, in fact, a critical aspect of the better-performing outcomes of AI tools, since this was reflected in standardized and specific prompts designed in this study. Specialized input was thus provided to enable them to outperform the results from previous performance benchmarks that did research pointing at such gaps in results. This flexibility further increased expectations that AI could find relevant uses in areas such as education and professional writing, requiring high accuracy and speed. At the same time, these divergent results with human ratings raised important questions about the context in which AI tools were used since human judgment was required for more subtle and creative assessments. Looking ahead, the integration of AI in writing tasks offered significant opportunities for improving learning outcomes and professional productivity. However, it also raised questions about what is the best way to balance AI capabilities with human insight. Future advancements in AI could focus on bridging this gap by incorporating more context-sensitive and content-driven feedback mechanisms, making these tools even more versatile. Ultimately, the present study confirmed that AI tools were not only valuable aids but also reshaping the way feedback was delivered, learned, and acted upon, paving the way for more synergistic applications of AI and human expertise in writing tasks. AI in writing held great potential for future learning and professional CONCLUSION AND SUGGESTIONS The aim of this study was to compare AI tools. Gemini and ChatGPT, and human experts in providing feedback on the employment application letters of non-native English-speaking students. Using a case study approach and six writing standards with a focus on grammar and mechanics, the study concluded AI technologyAiparticularly when used togetherAipossesses great potential in the quantity and accuracy of feedback. Overall. Gemini and ChatGPT Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati outperformed the two human raters in detecting grammatical and mechanical This was more noticeable in micro-level errors such as subject-verb agreement, spelling, capitalization, and punctuation. The ability to apply rules consistently, give feedback instantly, and check multiple linguistic features simultaneously gave AI a significant edge in error detection. Of the two. Gemini was more responsive to grammatical inconsistencies, such as subject-verb agreement and punctuation, but ChatGPT had greater functionality in generating variations of text and overall coherence. While these results showed AIAos potential. Human Experts displayed strength in identifying issues that required contextual awareness, such as tone appropriateness and subtle inconsistencies in style and content. This pointed to the necessity of human judgment, and also the rule-based character of AI tools. Notably. Gemini often provided suggestions for improvement without explicitly stating the errors, while ChatGPT provided high-level tips and even paraphrased chunks of content to recommend improvements towards proper writing. These strengths form the basis of the versatility of AI tools in learning contexts where students could benefit from error detection and guided improvement. Pedagogically, these findings suggested that AI could be a useful additional feedback provider, especially on repetitive and technical writing issues. However. AI was not a replacement for human instruction. Instead, an integrated feedback systemAiwhere AI offers rule-based feedback and human reviewers provide feedback on tone, audience sensitivity, and content coherenceAimight provide the most balanced and comprehensive writing support. Although this study employed a case study approach and did not seek broad generalisability, it opens several avenues for future research. One important direction involves refining AI tools to provide more context-sensitive feedback, such as adjustments in tone or appropriateness for a specific audienceAian area where human evaluators currently maintain an advantage. Additionally, understanding how students perceive and respond to feedback from artificial intelligence compared to human reviewers remains an open question, particularly in relation to long-term learning gains. Longitudinal studies that track writing development over time would be especially valuable in this regard. Another key area for investigation is how students receive and incorporate both human and AI-generated feedback, particularly when engaged in high-stakes writing tasks like job applications. Given that this study focused solely on job application letters, future research could also explore the effectiveness of AI feedback across different writing genresAisuch as argumentative essays, reports, or creative writingAiand among learners at various proficiency levels. Furthermore, as AI writing tools continue to evolve rapidly, comparative studies between platforms like Gemini. ChatGPT, and Claude could help identify relative strengths and inform best practices for educational implementation. Overall, the integration of AI software such as Gemini and ChatGPT in writing evaluation was transformative in potential. When combined with human Journal of English Language Teaching Innovations and Materials (Jelti. , 7. , 54-72 e-ISSN 2657-1617 Lutfi Arief Noerjaman. Ruswan Dallyono. Farida Hidayati judgment, such tools could shape a dynamic, equitable, and efficient feedback community that enhanced writing quality and learner agency. Rather than replacing human evaluators. AI acted as a catalyst that helped develop more scalable and efficient grading models. The potential of writing commentary lied in leveraging the complementarity between the accuracy of AI and the judgment of human beingsAia complementarity that held the promise of more nimble and integrated education for writing. REFERENCES