top of page

Welcome
to NumpyNinja Blogs

NumpyNinja: Blogs. Demystifying Tech,

One Blog at a Time.
Millions of views. 

Spring AI - Part 2: Dynamic Prompts, System Messages, and Streaming

Feb 21
5 min read

Welcome back! In Part 1, we built a simple HTTP endpoint that sends a user message to an AI model and returns plain text. In Part 2, we go further by writing cleaner code, building reusable prompts, controlling the AI's behaviour, and streaming responses in real time.

Source: NotebookLM
Source: NotebookLM

  1. What you'll build

By the end of this post, you will have:

  • A clean service layer (AiService) to keep AI logic separate from the controller.

  • A new endpoint that uses template variables (e.g., {topic} ) so you can reuse the prompt patterns easily.

  • A new endpoint that uses a system message to control how the model behaves (tone, style, rules).

  • A streaming endpoint using ChatClient.stream() that returns the answer progressively.


  1. Prerequisites

  • Your Part 1 project runs successfully with OpenAI or Ollama.

  • No new dependencies are needed for templates or system messages.

  • One new dependency is required for streaming (WebFlux, added in Section 6).


  1. Refactor: Move AI Logic into a Service

In Part 1, the AI call lived directly inside the controller. As we add more endpoints, it is better to move the logic into a @Service class. This keeps controllers focused on HTTP only.


Create src/main/java/com/spring/ai/springaidemo/AiService.java:

package com.spring.ai.spring_ai_demo;

import org.springframework.ai.chat.client.ChatClient;
import org.springframework.stereotype.Service;

@Service
public class AiService {

    private final ChatClient chatClient;

    public AiService(ChatClient.Builder builder) {
        // Spring Boot autoconfigures ChatClient.Builder for your chosen provider.
        // Calling builder.build() gives a ready-to-use ChatClient.
        this.chatClient = builder.build();
    }

    public String simpleAnswer(String message) {
        return chatClient
                .prompt()
                .user(message)
                .call()
                .content();
    }
}

Update AiController.java to use the service:

package com.spring.ai.spring_ai_demo;

import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.RequestParam;
import org.springframework.web.bind.annotation.RestController;

@RestController
public class AiController {

    private final AiService aiService;

    // Spring injects AiService automatically
    public AiController(AiService aiService) {
        this.aiService = aiService;
    }

    @GetMapping("/ai/hello")
    public String hello(@RequestParam(defaultValue = "Hello World!") String message) {
        return aiService.simpleAnswer(message);
    }
}

Test (same as Part 1):


  1. Understanding the Service Layer

4.1 @Service

@Service
public class AiService{
  • Marks the class as a Spring-managed service bean.

  • Spring creates one instance and injects it wherever needed (e.g., into AiController).

  • Keeps AI logic in one place - easy to update or test.


4.2 Constructor Injection

public AiService(ChatClient.Builder builder) {
	this.chatClient = builder.build();
}
  • ChatClient.Builder is auto-providedby Spring Boot based on your application.properties.

  • No manual HTTP configuration needed - Spring AI handles base URLs, API keys, and model selection


  1. Prompt Templates (Dynamic Prompts)

In Part 1, you passed raw user text straight to the model. Prompt templates let you write a reusable structure with {placeholder} variables filled at runtime.


5.1 Add the Method to AiService


Add this method to AiService.java:

public String explainInStyle(String topic, String audience) {
        return chatClient
                .prompt()
                .user(u -> u
                        .text("Explain \"{topic}\" to a {audience}. Keep it short and clear.")
                        .param("topic", topic)      // Replaces {topic} in the template
                        .param("audience", audience)) // Replaces {audience} in the template
                .call()
                .content();
}

5.2 Add the Endpoint to AiController


Add this method to AiController.java:

@GetMapping("/ai/explain")
public String explain(
		@RequestParam String topic,
		@RequestParam(defaultValue = "beginner") String audience) {
	return aiService.explainInStyle(topic, audience);
}

5.3 How it works

.user(u -> u
	.text("Explain \"{topic}\" to a {audience}.")
	.param("topic", topic)
	.param("audience", audience))
  • .text(...) defines the prompt pattern.

  • .param("key", value) safely substitutes {key}in the text.

  • No manual String.format() or concatenation needed - cleaner and safer.


Test URLs:

  • http://localhost:8080/ai/explain?topic=Spring AI&audience=beginner

  • http://localhost:8080/ai/explain?topic=Spring&audience=senior) developer


  1. System Messages (Control AI Tone and Rules)

AI prompts have roles: system, user, and assistant. The system role sets the behavior and tone of the model. The user role is what the end user asks.


6.1 Add Tutor Mode to AiService.java


Add this method to  AiService.java:

public String tutorAnswer(String question){
	return chatClient
          	.prompt()
             	.system("""
              	You are a friendly Java tutor. 
                 	- Use simple and clear language. 
                 	- Use short paragraphs. 
                 	- If you mention code, keep it minimal.
                	- Always encourage the learner. 
                	""")
          	.user(question)// The learner's actual question
             	.call()
             	.content();
}

6.2 Add Endpoint to AiController

@GetMapping("/ai/tutor")
public String tutor(@RequestParam String question){
	return aiService.tutorAnswer(question);
}

6.3 How it works

.system("You are a friendly Java tutor...")
.user(question)
  • system(...) runs before the user message and sets consistent rules for every call.

  • The model now always replies in a friendly, simple way - no need to repeat it in every request.


Test URLs:



  1. Streaming Responses (Real-Time Output)

So far, every endpoint waits for the complete response before returning it. Streaming changes this: it returns text chunk by chunk as the model generates it - exactly like you see in ChatGPT.


Spring AI supports streaming using stream() on the ChatClient, which returns a reactive Flux<String>.


7.1 Add WebFlux Dependency


Open pom.xml -> Add inside <dependencies>:

<dependency>
	<groupId>org.springframework.boot</groupId>
	<artifactId>spring-boot-starter-webflux</artifactId>
</dependency>	

Run: ./mvnw dependency:resolve


7.2 Add Streaming Method to AiService


Add these imports + method to AiService.java:

public Flux<String> streamAnswer(String message){
	return chatClient
		.prompt()
        	.user(message)
       	.stream()// Switch from .call() to .stream() for reactive output
       	.content(); // Returns Flux<String> - each chunk is a token from the model
}

7.3 Add Streaming Endpoint to AiController


Add these imports + method to AiController.java:

@GetMapping(value = "/ai/stream", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
public Flux<String> stream(@RequestParam String message){	// TEXT_EVENT_STREAM_VALUE sends chunks as Server-Sent Events (SSE)
	return aiService.streamAnswer(message);
}

7.4 How it works

.stream() // Instead of .call()
.content() // Returns Flux<String>
  • .call() -> Blocking: waits for full response, returns String.

  • stream() -> Recative: returns Flux<String>, each emission is a token/word chunk.

  • TEXT_EVENT_STREAM_VALUE sends each chunk over HTTP via Server-Sent Events (SSE).


Test in Browser:


Live streaming! Words appear progressively (like ChatGPT)
Live streaming! Words appear progressively (like ChatGPT)
  1. Final Complete Files

AiService.java

package com.spring.ai.spring_ai_demo;

import org.springframework.ai.chat.client.ChatClient;
import org.springframework.stereotype.Service;
import reactor.core.publisher.Flux;

@Service
public class AiService {

    private final ChatClient chatClient;

    public AiService(ChatClient.Builder builder) {
        this.chatClient = builder.build();
    }

    // Part 1: Simple Answer
    public String simpleAnswer(String message) {
        return chatClient
                .prompt()
                .user(message)
                .call()
                .content();
    }

    // Part 2a: Prompt templates
    public String explainInStyle(String topic, String audience) {
        return chatClient
                .prompt()
                .user(u -> u
                        .text("Explain \"{topic}\" to a {audience}. Keep it short and clear.")
                        .param("topic", topic)     
                        .param("audience", audience))
                .call()
                .content();
    }

    // Part 2b: System message
    public String tutorAnswer(String question){
        return chatClient
                .prompt()
                .system("""
                        You are a friendly Java tutor. 
                        - Use simple and clear language. 
                        - Use short paragraphs. 
                        - If you mention code, keep it minimal.
                        - Always encourage the learner. 
                        """)
                .user(question)   
                .call()
                .content();
    }

    // Part 2c: Streaming
    public Flux<String> streamAnswer(String message){
        return chatClient
                .prompt()
                .user(message)
                .stream()
                .content(); 
    }
}

AiController.java

package com.spring.ai.spring_ai_demo;

import org.springframework.http.MediaType;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.RequestParam;
import org.springframework.web.bind.annotation.RestController;
import reactor.core.publisher.Flux;

import java.awt.*;

@RestController
public class AiController {

    private final AiService aiService;
    
    public AiController(AiService aiService) {
        this.aiService = aiService;
    }

    @GetMapping("/ai/hello")
    public String hello(@RequestParam(defaultValue = "Hello World!") String message) {
        return aiService.simpleAnswer(message);
    }

    @GetMapping("/ai/explain")
    public String explain(
            @RequestParam String topic,
            @RequestParam(defaultValue = "beginner") String audience) {
        return aiService.explainInStyle(topic, audience);
    }

    @GetMapping("ai/tutor")
    public String tutor(@RequestParam String question){
        return aiService.tutorAnswer(question);
    }

    @GetMapping(value = "ai/stream", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
    public Flux<String> stream(@RequestParam String message){
        return aiService.streamAnswer(message);
    }
}
  1. What You Have Achieved

At this point, you have:

  • A service layer keeping AI logic clean and reusable.

  • Prompt templates with {placeholders} for dynamic, reusable prompts.

  • System messages for consistent AI behaviour across all requests.

  • Streaming real-time responses using Flux<String> and SSE.

In Part 3, we will add:

  • Structured output - map AI responses directly into Java records/POJOs

  • Validation and error handling - gracefully handle bad inputs and API errors.

  • Multiple system personas - switch between tutor, code reviewer, and summarizer roles.


You've done it! From raw prompts to a full AI service layer.

Questions? Comment your test results below.


Source: Spring AI Reference Docs - ChatClient API, Prompt Templates



 
 

+1 (302) 200-8320

NumPy_Ninja_Logo (1).png

Numpy Ninja Inc. 8 The Grn Ste A Dover, DE 19901

© Copyright 2025 by Numpy Ninja Inc.

  • Twitter
  • LinkedIn
bottom of page