top of page

Welcome
to NumpyNinja Blogs

NumpyNinja: Blogs. Demystifying Tech,

One Blog at a Time.
Millions of views. 

What Is Tesseract and how it is used in test automation world?

Apr 30, 2025
4 min read

Updated: May 1, 2025

Hi everyone. Hope you are doing well. Today i will walk you through all about tesseract.

What is tesseract, how it is used in java and the implementation of it in test automation framework.

If you want to know the definition of tesseract, then it is an open-source optical character

recognition (OCR) engine developed by HP and maintained by Google that is used to extract text from images .

Tesseract can recognize and extract text from images in over 100 languages. Basically in selenium or any other Dom based tools, we can not validate CAPTCHA images, text rendered in images, PDFs, Image-based buttons, Charts and reports rendered as images, scanned documents, or non-HTML content . So to automate these features, this tesseract comes into picture in test automation world. Lets go and dig into it.



This is the flow how text can be extracted from image
This is the flow how text can be extracted from image

These are the Prerequisites needed to implement in testing

  • Java (jdk)

  • maven based project

  • Selenium WebDriver

  • Cucumber

  • Tesseract software need to install locally

  • tess4j library


Step1:  First Install tesseract for windows.

We can download it from below link

Step:2

For implementation, we need to add tess4j Library for java project and in

python, need to add pytesseract library which provides an interface to tessract's OCR engine.

Tess4J

Now lets know about tess4j : Tess4J is a Java wrapper for the Tesseract APIs that provides OCR support for various image formats like JPEG, GIF, PNG, and BMP. It is written in C++.So we need a java wrapper which is tess4j acts like a bridge for seamless integration with Java projects.


So internally how it works. Actually it uses JNI-Java native interface to call functions from tesseract C++ library.

There is one method named as tesseract.doOCR(image file),which is the core function to extract text from image.

when we run this method ,

  • It loads the image using Java's BufferedImage.

  • sends it to the native Tesseract engine

  • extracts and returns text as a string to Java.


There are 3 major classes which are used in Tess4J library.

Tesseract- Main class to perform OCR.

TesseractException - It handles OCR-related errors

ITesseract -Interface for the Tesseract class.


eng.traineddata file

We should also know about eng.traineddata file as part of tesseract implementation. It is a language data file used by the Tesseract OCR engine to recognize English text in images.

  • It contains neural network models and rules that Tesseract uses to detect and decode characters and words for the English language.

  • It is part of the Tesseract "tessdata" folder.


tesseract.setLanguage("eng"); // use eng.traineddata - we need to write this code to set language as english.

we can download more language files from the official repo:


Lets see the implementation

This maven dependency need to be added for tess4j to implement.

<dependency>

    <groupId>net.sourceforge.tess4j</groupId>

    <artifactId>tess4j</artifactId>

    <version>5.15.0</version>

</dependency>


The latest maven dependency need to be added in pom.xml file


Lets see the below use case for tesseract .So below is an image, We need to extract the text "LMS-learning Management System " from this image.


So to start over, We have to see the project structure. Below is the diagram of folder structure.

So in feature file we can write the test case as a Scenario with Given, When and Then format.

To implement the above use case the scenario will be like below:


In feature file  ,

 Scenario: Verify application name 

   Given Admin is in landing page

   When Admin gives the correct portal URL

  Then Admin should see  LMS - Learning Management System


In Step definition file

@Then("Admin should see LMS - Learning Management System")

public void admin_should_see_lms_learning_management_system() throws TesseractException, IOException {

String LMS_text=loginpg.image_to_text_converter();

System.out.println(loginpg.image_to_text_converter());
if(LMS_text.contains("LMS - Learning Management System")) {
Assert.assertTrue(true);
}
else {
Assert.assertTrue(false);
}
}

So the above code states that we are calling image_to_text_converter() which explains about image extraction from an image file and then we are asserting the text from image to text of "LMS - Learning Management System"


Then lets see the working code of tesseract which I coded inside  image_to_text_converter() method.

This method I declared inside PageObject package.

public String image_to_text_converter() throws TesseractException, IOException{

WebElement imageElement = driver.findElement(By.cssSelector("img.images"));

// Capture the image as a file

File imageFile = imageElement.getScreenshotAs(OutputType.FILE);

// Define the destination file path

File savedImage = new File("./src/test/resources/testData/image.png");

ImageIO.write(ImageIO.read(imageFile), "png", savedImage);

// Set up Tesseract OCR

Tesseract tesseract = new Tesseract();

tesseract.setDatapath("C:\\Program Files\\Tesseract-OCR\\tessdata");

// use eng.traineddata  to set the language as english

tesseract.setLanguage("eng"); 

// Extract text from the image

String extractedText = tesseract.doOCR(savedImage);

System.out.println("Extracted Text: " + extractedText);

// Close the browser

return extractedText;

}

Lets understand the code line by line. So public String image_to_text_converter() throws TesseractException, IOException method Captures a screenshot of an image element from a webpage using Selenium. Saves the image locally. Used Tess4J to extract text from the saved image .Returns the extracted text.

Tesseract tesseract = new Tesseract();
tesseract.setDatapath("C:\\Program Files\\Tesseract-OCR\\tessdata");

This piece of code Creates a Tesseract instance.

Sets the data path to the Tesseract installation folder where eng.traineddata and other language files are located.


String extractedText = tesseract.doOCR(savedImage);

This Passes the saved image file to doOCR().Tesseract scans the image and returns extracted text.


System.out.println("Extracted Text: " + extractedText);
return extractedText;

This logs the result to the console and returns the text to be used in assertions and in reports.



This is the implementation of tesseract
This is the implementation of tesseract

Hope this blog will make you understand all about Tesseract and its implementation. “That’s it for today — stay curious, stay kind, and I’ll catch you in the next one.” Thank you.


Happy learning. Cheers!!!!!!.


 
 

+1 (302) 200-8320

NumPy_Ninja_Logo (1).png

Numpy Ninja Inc. 8 The Grn Ste A Dover, DE 19901

© Copyright 2025 by Numpy Ninja Inc.

  • Twitter
  • LinkedIn
bottom of page