Text extracted From an Image in Selenium Automation:
Updated: Feb 15, 2025
Steps to Extract Text from an Image in Selenium with Java:
1. Capture the Image Element: Use Selenium to locate the image on the webpage and take a screenshot.
2. Save the Screenshot: Save the screenshot of the image to your local system.
3. Apply OCR: Use Tesseract OCR to extract text from the saved image.
Key Points:
Install Dependencies:
1. Selenium WebDriver for Java.
2. Tess4J for OCR integration with Java.
Set Tesseract Data Path: Ensure Tesseract's test data directory is properly set in your environment or explicitly in your code
l Handle Image Quality: OCR accuracy depends on the quality and clarity of the image. Pre-processing the image may help improve results.
Limitations:
Complex images with distortions or overlapping text may yield inaccurate results.
Ensure that the Tesseract language pack for the text in the image is installed.
To extract text from an image, first install Tesseract, add dependencies, and write a program to save an image and retrieve text from an image.
1. Go to the Tesseract Releases page and download the lastest .exeinstaller for Windows. Run the installer. Choose the directory to install Tesseract. Better choose the “c:\Program Files\Tesseract-OCR” directory in your system. During installation, select “Additional Language Data” and “Script Data” . After installation, Go to Control Panel> System>Advanced System Settings>Environment Variables
Under “System variables” find the path. click Edit, and add “c:\Program Files\Tesseract-OCR”.To verify the installation, open the command prompt and type tesseract -v, you should see the tesseract version and installation confirmation.

Here, we are using BDD framework In that pom.xml, add tess4j and java-ocr-api dependencies for extract text from an image. In pageObject class, the script was in the method image_to_text_converter() and called this method in StepDefination where steps from featurefile Scenario's will map to the StepDefination.
1. Add dependencies, tess4j, and java-ocr-api.

Tess4J is a Java wrapper for the Tesseract OCR (Optical Character Recognition) engine. It allows developers to integrate Tesseract’s OCR capabilities into Java applications, it is easy to extract text from images or scanned documents pro-grammatically.
Tess4J used for: Extracts text from scanned documents, PDF s,or images.Automating data entry from physical documents.Reading text from photographs, receipts, invoices, or IDs.Extracting text from CAPTCHA images for automated workflows. Converts printed text into digital text from screen readers
Java-ocr-api: The java-ocr-api, also known as the Asprise Java OCR SDK, is a library used to perform Optical Character Recognition (OCR) and barcode recognition in Java applications.
Core Functionality:
Text Extraction: Convert images (JPEG, PNG, TIFF, PDF, etc.) into machine-readable text.
Barcode Recognition: Read various barcode formats like Code 128, EAN, UPC, QR codes, and more.
Searchable PDF Creation: Generate searchable PDF files from scanned documents.
Data Capture: Extract structured data from tables and forms.
Multi-Language Support: Recognize text in over 20 languages, including English, Spanish, French, German, and more.
1. Write a program for extracting text from an Image :
Here, In the Page Object class, I used the cssSelector locator to find an Image Element. I am capturing an image as a file using ‘image element.getScreenshotAs(OUTPUTTYPE.FILE)’ in File. I am pointing to the an image to save in a specified location with the ‘.png’ extension.
Now set up Tesseract OCR using the line ‘Tesseract tesseract=new Tesseract();’.Set the Tesseract Data path where my Tesseract-Ocr is located using the line 'tesseract.setDATApath(“C;\\Program Files\\Tesseract-OCR\\tessdata”);'.
Now extracted text from an image that is saved in File using the line 'tesseract.doOCR (savedImage);'. Now print the extracted text.


From StepDefinition class, using loginPage object, the script called image_to_text_converter method, and stored that text in numpyninja_text using the line 'String numpyninja_text=loginPage.image_to_text_converter()' and find that numpyninja_text contains our specified text using line 'if(numpyninja_text.contains("NumpyNinja"))' and validate that text contains in the text which retrieved text from an image using assertions.


