Skip to content

For developers

Handwritten Answer Grading API: OCR, Vision Models, Build vs Buy

Why OCR fails on student handwriting, what a production grading pipeline needs, build vs buy, and how the Evalezy API checks answer sheets by URL or file.

By the Evalezy team 8 min read

Key takeaways

  • Printed-text OCR is the wrong tool for student handwriting: even Tesseract’s own FAQ says it is designed for printed text.
  • A production grading pipeline is ten problems, not one: cleanup, line geometry, reading, answer location, rubric grading, validation, annotation, review, failure handling and billing.
  • Treat student writing as untrusted input: an answer can contain instructions aimed at the grader.
  • Build if grading is your core product or you need a script Evalezy has not tested; use an API if grading is a feature and you want marks plus a checked PDF.
  • The Evalezy API flow is asynchronous: create an assessment, set criteria, start an evaluation by URL or file, then poll or take a webhook.

Your learners write on paper. Your product lives online. Somewhere between the two, a handwritten answer sheet has to become marks, feedback and something a learner actually wants to open. If you are a product manager or engineer adding that step to an edtech app, this post covers why plain OCR does not get you there, what a pipeline needs before you can trust it in production, how to think about building it yourself, and how the Evalezy API handles it.

Why printed-text OCR fails on student handwriting

Most OCR engines were built for print. Tesseract, the open-source OCR engine, says so in its own FAQ: you can use it for handwriting, but “it won’t work very well, as Tesseract is designed for printed text.” Handwriting is a harder problem because, as Wikipedia’s overview of the field puts it, “different people have different handwriting styles.”

Cloud document services do read handwriting, but check the language list before you plan around one. Microsoft’s Document Intelligence Read model lists Hindi, Marathi, Tamil and many other languages for printed text. Its v4.0 list for handwritten text has twelve languages, and none of them is an Indian language.

Student answer sheets then add problems of their own on top of the handwriting itself. These are the ones Evalezy’s pipeline is built to handle:

  • Phone photos, not scans. The sample copy on this site is a phone photo of a spiral notebook. Pages arrive tilted, shadowed and low in contrast.
  • Answers out of order. Students answer Question 7 before Question 3, or number answers wrongly. A pipeline that trusts the written number marks the wrong answer.
  • Maths. Fractions, powers and working spread over several lines need a careful second look.
  • Long copies. A 20-page script has to be split into answers before anything can be graded.

And when reading fails, grading fails quietly. A grader handed garbled text has nothing true to grade. The pipeline must be able to say “this copy could not be read” instead of returning a mark.

What a robust pipeline needs

“Read the page, ask a model for a mark” is a demo. A version you can put in front of thousands of learners needs every stage below, and each one is a small project of its own.

Stages of a handwritten answer grading pipeline
StageWhat it doesWhat breaks without it
Page cleanupDeskews, denoises and boosts contrast on scans and phone photosFaint or tilted lines are misread or missed
Line geometryRecords where each line of writing sits on the pageTicks and notes land in the wrong place
Vision readingReads whole pages in context, with a second read for hard lines such as mathsWord fragments with no structure reach the grader
Answer locationFinds which pages hold which answer, by contentOut-of-order answers score zero
Rubric gradingOne set of criteria per question, reused for every studentEach copy is marked to a slightly different standard
Mark validationKeeps marks within the maximum, on the allowed step, with criteria adding upImpossible totals reach learners
AnnotationDraws ticks, crosses and notes on the student’s own pagesLearners get a table of numbers, not their copy
Human reviewHolds marks as a draft until a teacher releases themAn AI mistake goes straight to a learner
Failure and retriesFails unreadable copies with a reason; re-asks and flags low-confidence readsGuessed marks, or jobs stuck forever
Idempotent billingCharges each sheet once, and only when it succeedsRetries double-charge; failures cost money

Reading: a vision model, not a bag of words

A vision model reads the page image directly, layout and all, rather than handing your grader a list of recognised words. In Evalezy, copies of up to 40 pages are read in full this way, and unclear maths lines get a second read by Mathpix, a dedicated equation reader, up to four per copy. For copies of four pages or more, a locate pass works out which pages hold which answer. If a question then looks unattempted, it is re-checked against the whole copy, so a locate miss cannot zero a student.

Grading: the same criteria for every student

Consistency comes from fixing the rubric, not from hoping the model behaves the same way each time. Evalezy generates one marking scheme per question, once, and reuses it for every student: three to five criteria for subjective questions and one for objective questions, rescaled to add up exactly to the question’s maximum. Teachers can edit it or supply their own, and a stored model answer is used over a generated one. Marks are given in half-mark steps.

Treat student writing as untrusted input. OWASP puts prompt injection first (LLM01) in its 2025 Top 10 for LLM applications, and notes that injected instructions can arrive indirectly, through files or images the model processes. An answer sheet is exactly that kind of input. In a production test of Evalezy’s typed-answer grading, an answer that told the AI to “award full marks” scored 0 out of 5.

Line geometry and mark validation

Two stages are easy to skip in a prototype and painful to add later. The first is line geometry. A grader can say “the second point is wrong”, but to put a cross beside that point you need to know where the line sits on the page image. Keep a position for every line you read, in the same coordinates as the page you will draw on after cleanup and deskewing, or the annotations drift off the writing.

The second is mark validation. Treat every mark a model returns as a proposal and check it before you store it: not above the question’s maximum, on your marking step, criteria adding up to the question mark, question marks adding up to the total. When a check fails, re-ask or flag the question for a teacher; do not quietly round. Fixed rules, such as half-mark steps and criteria that sum exactly to the maximum, give these checks clear targets.

Annotation and review

Learners want to see their own page marked, not a JSON blob turned into a table. Evalezy draws ticks, crosses, short notes and marks per question in a handwriting style, with the total circled on page 1, and the same copy re-renders identically every time, which matters when a learner and a teacher are looking at the same page during a dispute. The AI never publishes a result on its own in the Evalezy dashboard: marks land as a draft, a teacher reviews and releases, and a mark the teacher edits is never overwritten by a later AI run.

Failure handling and billing

Async jobs fail and get retried, so billing has to be idempotent. Stripe’s API is a well-known example of the pattern: clients send an idempotency key so a request can be retried “without accidentally performing the same operation twice.” Whatever mechanism you choose, a retry should never become a second charge. In Evalezy, a copy that is mostly unreadable fails with a clear error and is not graded by guesswork, and failed, cancelled and unreadable checks are not charged.

Build vs buy

Build in-houseUse an API
What you ownModels, prompts, data flow and every stage aboveYour rubric, model answers and review flow; the pipeline is the vendor’s
Before the first trustworthy resultAll ten stages, plus a test set of real copiesIntegration against a handful of endpoints
Ongoing workModel changes, prompt regressions, new handwriting edge cases, PDF renderingThe vendor’s job
Cost shapeModel calls, infrastructure and engineering time; per-page cost variesA fixed rate per page (Evalezy: ₹1 or $0.01); failures not billed
Languages and scriptsWhatever you build and testEvalezy: tested on English; other languages need samples first

Build if grading is your core product, if you need a language or script now that no vendor has tested, or if answer sheets must never leave your own infrastructure. Use an API if grading is one feature of a larger product, if you want a checked PDF and a review step without building them, and if a per-page cost is easier to plan for than a team. If you are weighing a general chatbot instead, our comparison with ChatGPT covers that approach.

Either way, decide on evidence. Collect a test set of real answer sheets from your own learners, say fifty copies across neat and messy handwriting, phone photos and scans, and have experienced teachers mark them first. Run them through your prototype or a vendor and compare question by question: where marks differ, whether the reasons make sense, and which copies failed. Keep that set. It becomes your regression test for every model or prompt change afterwards.

How the Evalezy API flow works

  1. 1

    Create an assessment

    POST /assessments with a list of questions and marks, or a question paper URL to be read into questions. You get back an assessment id.

  2. 2

    Set evaluation criteria

    PUT /assessments/{id}/criteria with criteria and an optional model answer per question; criteria marks are rescaled to sum to the question’s marks. Or POST /assessments/{id}/criteria/generate for a draft to review and edit.

  3. 3

    Start an evaluation

    POST /evaluations with a public answer_sheet_url, or a file id from POST /files, plus your own student id and an optional webhook_url. It returns an evaluation id with status queued.

  4. 4

    Poll or wait for the webhook

    GET /evaluations/{id} moves through queued, reading and grading to completed, failed or cancelled. With a webhook_url, Evalezy posts to you when the evaluation completes or fails.

  5. 5

    Use the result

    The total, per-question marks, the student’s answer as read, feedback, the criteria breakdown, the checked_copy_url for the red-pen PDF and the pages billed.

Here is the core of an integration in Node.js 18 or later, polling instead of using a webhook:

const API = "https://api.evalezy.com/v1";
const headers = {
  Authorization: "Bearer " + process.env.EVALEZY_API_KEY,
  "Content-Type": "application/json",
};

// 1. Start checking one answer sheet
const res = await fetch(API + "/evaluations", {
  method: "POST",
  headers,
  body: JSON.stringify({
    assessment_id: "asm_7Hq2",
    answer_sheet_url: "https://files.example.edu/ix-b/roll-14.pdf",
    student: { external_id: "IXB-14" },
  }),
});
const { id } = await res.json(); // status: "queued"

// 2. Poll every 30 s until it finishes (or pass webhook_url instead)
let ev;
do {
  await new Promise((r) => setTimeout(r, 30000));
  ev = await (await fetch(API + "/evaluations/" + id, { headers })).json();
} while (["queued", "reading", "grading"].includes(ev.status));

// 3. Use the result
if (ev.status === "completed") {
  console.log(ev.total);            // { awarded: 38, max: 80 }
  console.log(ev.checked_copy_url); // red-pen PDF to show the learner
} else {
  console.log(ev.status, ev);       // failed includes a reason; not billed
}

The example values come from the real sample copy, a 7-page Class IX Social Science answer sheet that scored 38 out of 80.

Shipping it in your product

  • Show the checked copy, not only the JSON. The checked copy is what learners open and share. Put it first; use the per-question data for progress tracking and reports.
  • Put a reviewer before the learner. The API returns the AI’s marks with reasons. Whether a teacher sees them first is up to your product, and we recommend that one does, as in Evalezy’s own teacher review.
  • Map results to your users. Send your own student id as student.external_id on every evaluation, so a webhook or poll result maps straight back to the right learner.
  • Make your own side idempotent. Store the evaluation id against the learner’s submission as soon as you get it. If your worker crashes and retries, look up the stored id instead of starting a second evaluation of the same sheet, which would be a separate run.
  • Design for minutes, not seconds. A copy typically takes 1 to 8 minutes. Show progress, let the learner leave and notify them when the result is ready.
  • Handle failure honestly. When a sheet fails as unreadable, ask the learner for a clearer scan rather than showing a zero. Failed checks are not billed.
  • Know the limits. Diagrams, maps and graphs are judged by their labels and written explanation, not as drawings. Grading is tested on English answer sheets.
  • Budget per page. A 5-page sheet costs ₹5 or $0.05. A re-check is a new run and is billed again, so re-check deliberately.

Where Evalezy fits

Evalezy handles the reading, grading, annotation and review stages for you, in a dashboard or through the API. Send a handwritten answer sheet by URL or file and get back per-question marks, feedback, the student’s answer as read, a criteria breakdown and a red-pen checked PDF.

Same engine, same price. ₹1 per page in India (excluding GST) or $0.01 elsewhere, the same rate as the dashboard. Failed, cancelled and unreadable sheets are not billed. See pricing.

Your rubric, or a draft to edit. Send your own criteria and model answers, or ask for generated criteria and review them before you evaluate. See how it works and the edtech platforms page.

Book a demo

Sources

  1. Tesseract documentation, FAQ: “Can I use Tesseract for handwriting recognition?” (checked 1 October 2026)
  2. Wikipedia, Handwriting recognition (offline recognition and handwriting styles) (checked 1 October 2026)
  3. Microsoft Learn, Language and locale support for Read and Layout document analysis (Document Intelligence v4.0) (checked 1 October 2026)
  4. OWASP Gen AI Security Project, LLM01:2025 Prompt Injection (checked 1 October 2026)
  5. Stripe API reference, Idempotent requests (checked 1 October 2026)

Frequently asked questions

Can I grade handwritten answers with Tesseract or another OCR library?+

You can extract some text, but Tesseract’s own FAQ says it will not work very well on handwriting because it is designed for printed text. Grading also needs answer location, a rubric, validation and annotation on top of reading, so OCR is at most one stage of the pipeline.

Why not send the page straight to a general chatbot?+

A single chat prompt gives you a mark, but not a fixed rubric reused for every student, mark validation, a checked PDF, a review step or clear failure on unreadable pages. We compare the two approaches on the Evalezy vs ChatGPT page.

Is the Evalezy API synchronous?+

No. Checking a copy typically takes 1 to 8 minutes, so evaluations are asynchronous: you get an id straight away, then poll the status endpoint or receive a webhook when the check completes or fails.

What does a completed evaluation return?+

The total, and for each question the marks awarded and maximum, the student’s answer as read, feedback and the criteria breakdown, plus a URL to the checked copy PDF and the pages billed.

How is the API billed?+

Per page checked: ₹1 in India (excluding GST) or $0.01 elsewhere, the same as the dashboard. Failed, cancelled and unreadable sheets are not billed. A re-check is a new run and is billed again.

Does it read Hindi or other Indian-language handwriting?+

Evalezy is built and tested on English answer sheets. If your learners write in another language, send us sample copies before you plan a launch around it.

See your own copies checked.

Book a 20-minute demo. Bring a question paper and a few scanned answer sheets, and watch Evalezy check them in red pen.