Skip to content
elephantoo

The Streams API

Lesson 31 of 43 20 min read

Build pipelines with filter, map, sorted and reduce, then collect with groupingBy, joining and more.


The Streams API (java.util.stream, Java 8+) lets you process collections declaratively: say what you want ("the names of active users over 18, sorted"), not how to loop. Combined with lambdas, streams make data processing shorter, clearer and easy to parallelise.

Loop vs stream#

LoopVsStream.java
import java.util.ArrayList;
import java.util.Collections;
import java.util.List;

public class LoopVsStream {
    record User(String name, int age, boolean active) { }

    public static void main(String[] args) {
        List<User> users = List.of(
            new User("Ravi", 34, true), new User("Anu", 17, true),
            new User("Meera", 29, false), new User("Dev", 41, true));

        // Imperative: how to do it
        List<String> result = new ArrayList<>();
        for (User u : users) {
            if (u.active() && u.age() >= 18) {
                result.add(u.name().toUpperCase());
            }
        }
        Collections.sort(result);
        System.out.println(result);

        // Declarative: what we want
        List<String> names = users.stream()
                .filter(u -> u.active() && u.age() >= 18)
                .map(u -> u.name().toUpperCase())
                .sorted()
                .toList();
        System.out.println(names);
    }
}
Output
[DEV, RAVI]
[DEV, RAVI]

Anatomy of a stream pipeline#

Output
source  →  intermediate operations (0 or more)  →  terminal operation (exactly 1)
list.stream()  .filter(...)  .map(...)  .sorted()     .toList()
  1. Source: a collection (list.stream()), an array (Arrays.stream(arr)), Stream.of(...), a range (IntStream.range), a file (Files.lines)...
  2. Intermediate operations return a new stream: filter, map, sorted, distinct, limit... They are lazy: nothing happens yet.
  3. Terminal operation produces a result or side effect: toList, collect, forEach, count, reduce, findFirst... This triggers the processing.

Important properties:

  • A stream doesn't modify its source.
  • A stream can be consumed only once.
  • Elements flow through the pipeline one at a time, so with limit or findFirst, much of the source may never be processed.
Laziness.java
import java.util.List;

public class Laziness {
    public static void main(String[] args) {
        List<Integer> result = List.of(1, 2, 3, 4, 5, 6).stream()
                .filter(n -> { System.out.println("filter " + n); return n % 2 == 0; })
                .map(n -> { System.out.println("map " + n); return n * n; })
                .limit(2)
                .toList();
        System.out.println(result);
    }
}
Output
filter 1
filter 2
map 2
filter 3
filter 4
map 4
[4, 16]

Elements 5 and 6 are never touched: once limit(2) is satisfied, the pipeline stops.

Core intermediate operations#

Intermediate.java
import java.util.Comparator;
import java.util.List;
import java.util.stream.Stream;

public class Intermediate {
    public static void main(String[] args) {
        List<String> words = List.of("banana", "Apple", "cherry", "apple", "date", "banana");

        System.out.println(words.stream().filter(w -> w.length() > 5).toList());
        System.out.println(words.stream().map(String::length).toList());
        System.out.println(words.stream().map(String::toLowerCase).distinct().toList());
        System.out.println(words.stream().sorted().toList());
        System.out.println(words.stream().sorted(Comparator.comparing(String::length).reversed()).limit(2).toList());
        System.out.println(words.stream().skip(4).toList());

        List<List<Integer>> nested = List.of(List.of(1, 2), List.of(3), List.of(4, 5));
        System.out.println(nested.stream().flatMap(List::stream).toList());   // flatten

        System.out.println(Stream.of(3, 1, 4, 1, 5, 9, 2).takeWhile(n -> n < 5).toList());
        System.out.println(Stream.of(3, 1, 4, 1, 5, 9, 2).dropWhile(n -> n < 5).toList());
    }
}
Output
[banana, cherry, banana]
[6, 5, 6, 5, 4, 6]
[banana, apple, cherry, date]
[Apple, apple, banana, banana, cherry, date]
[banana, cherry]
[date, banana]
[1, 2, 3, 4, 5]
[3, 1, 4, 1]
[5, 9, 2]
OperationDoes
filter(predicate)keep matching elements
map(function)transform each element
flatMap(fn)map each element to a stream and flatten the results
distinct()remove duplicates (uses equals)
sorted() / sorted(cmp)sort
limit(n) / skip(n)take / drop the first n
takeWhile / dropWhiletake / drop while a condition holds (Java 9+)
peek(action)look at elements as they pass (for debugging)
mapToInt, mapToObj, boxedswitch between object and primitive streams

Terminal operations#

Terminal.java
import java.util.List;
import java.util.Optional;

public class Terminal {
    public static void main(String[] args) {
        List<Integer> nums = List.of(4, 8, 15, 16, 23, 42);

        System.out.println(nums.stream().count());
        System.out.println(nums.stream().anyMatch(n -> n > 40));
        System.out.println(nums.stream().allMatch(n -> n > 0));
        System.out.println(nums.stream().noneMatch(n -> n % 7 == 0));

        Optional<Integer> firstOdd = nums.stream().filter(n -> n % 2 == 1).findFirst();
        System.out.println(firstOdd.orElse(-1));

        System.out.println(nums.stream().max(Integer::compare).get());
        System.out.println(nums.stream().reduce(0, Integer::sum));         // fold into one value
        System.out.println(nums.stream().map(String::valueOf).reduce("", String::concat));

        nums.stream().filter(n -> n > 20).forEach(n -> System.out.print(n + " "));
        System.out.println();
    }
}
Output
6
true
true
false
15
42
108
4815162342
23 42

findFirst, max and min return an Optional, because the stream might be empty (next lesson). reduce(identity, accumulator) combines all elements into one value.

Primitive streams: IntStream, LongStream, DoubleStream#

For numbers, primitive streams avoid boxing and add handy maths methods:

Primitives.java
import java.util.IntSummaryStatistics;
import java.util.List;
import java.util.stream.IntStream;

public class Primitives {
    record Order(String id, int amount) { }

    public static void main(String[] args) {
        System.out.println(IntStream.rangeClosed(1, 100).sum());           // 1..100
        System.out.println(IntStream.range(0, 5).boxed().toList());        // 0..4

        List<Order> orders = List.of(new Order("A", 250), new Order("B", 900), new Order("C", 410));
        int total = orders.stream().mapToInt(Order::amount).sum();
        double avg = orders.stream().mapToInt(Order::amount).average().orElse(0);
        System.out.println(total + " " + avg);

        IntSummaryStatistics stats = orders.stream().mapToInt(Order::amount).summaryStatistics();
        System.out.println(stats.getMin() + " " + stats.getMax() + " " + stats.getCount());

        String squares = IntStream.rangeClosed(1, 5).map(i -> i * i)
                .mapToObj(String::valueOf).reduce((a, b) -> a + "," + b).orElse("");
        System.out.println(squares);
        System.out.println("hello".chars().filter(c -> c == 'l').count());
    }
}
Output
5050
[0, 1, 2, 3, 4]
1560 520.0
250 900 3
1,4,9,16,25
2

Collectors: building results#

collect(...) with the Collectors utility class builds lists, sets, maps, strings and statistics:

CollectorsDemo.java
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.TreeMap;
import java.util.stream.Collectors;

public class CollectorsDemo {
    record Employee(String name, String dept, double salary) { }

    public static void main(String[] args) {
        List<Employee> staff = List.of(
            new Employee("Asha", "Eng", 120_000), new Employee("Ben", "Eng", 95_000),
            new Employee("Chen", "Sales", 70_000), new Employee("Dia", "HR", 65_000),
            new Employee("Eli", "Sales", 82_000));

        List<String> names = staff.stream().map(Employee::name).collect(Collectors.toList());
        Set<String> depts = staff.stream().map(Employee::dept).collect(Collectors.toSet());
        String csv = staff.stream().map(Employee::name).collect(Collectors.joining(", ", "[", "]"));
        System.out.println(names.size() + " " + depts.size() + " " + csv);

        Map<String, Double> salaryByName = staff.stream()
                .collect(Collectors.toMap(Employee::name, Employee::salary));
        System.out.println(salaryByName.get("Chen"));

        Map<String, List<String>> byDept = staff.stream().collect(Collectors.groupingBy(
                Employee::dept, TreeMap::new, Collectors.mapping(Employee::name, Collectors.toList())));
        System.out.println(byDept);

        Map<String, Long> headcount = staff.stream()
                .collect(Collectors.groupingBy(Employee::dept, TreeMap::new, Collectors.counting()));
        System.out.println(headcount);

        Map<String, Double> avgSalary = staff.stream()
                .collect(Collectors.groupingBy(Employee::dept, TreeMap::new, Collectors.averagingDouble(Employee::salary)));
        System.out.println(avgSalary);

        Map<Boolean, List<String>> highEarners = staff.stream().collect(Collectors.partitioningBy(
                e -> e.salary() >= 90_000, Collectors.mapping(Employee::name, Collectors.toList())));
        System.out.println(highEarners);
    }
}
Output
5 3 [Asha, Ben, Chen, Dia, Eli]
70000.0
{Eng=[Asha, Ben], HR=[Dia], Sales=[Chen, Eli]}
{Eng=2, HR=1, Sales=2}
{Eng=107500.0, HR=65000.0, Sales=76000.0}
{false=[Chen, Dia, Eli], true=[Asha, Ben]}
CollectorResult
toList(), toSet()a List / Set
toMap(keyFn, valueFn)a Map (throws on duplicate keys unless you pass a merge function)
joining(sep, prefix, suffix)one String
groupingBy(classifier[, mapFactory][, downstream])Map<K, List<T>> or Map<K, downstream result>
partitioningBy(predicate[, downstream])Map<Boolean, ...>
counting(), summingInt, averagingDouble, maxBy, minByaggregates, usually as downstream collectors
mapping(fn, downstream)transform before collecting

stream.toList() (Java 16+) returns an unmodifiable list and is the most concise option. collect(Collectors.toList()) returns a mutable ArrayList in practice, but that is not guaranteed.

toMap with duplicate keys throws IllegalStateException. Supply a merge function to resolve them, e.g. Collectors.toMap(Employee::dept, Employee::salary, Double::sum).

A worked example: sales report#

SalesReport.java
import java.util.Comparator;
import java.util.List;
import java.util.Map;
import java.util.stream.Collectors;

public class SalesReport {
    record Sale(String region, String product, int units, double price) {
        double revenue() { return units * price; }
    }

    public static void main(String[] args) {
        List<Sale> sales = List.of(
            new Sale("North", "Laptop", 5, 55_000), new Sale("South", "Phone", 20, 18_000),
            new Sale("North", "Phone", 12, 18_000), new Sale("West", "Laptop", 2, 55_000),
            new Sale("South", "Tablet", 7, 25_000), new Sale("West", "Phone", 9, 18_000));

        double total = sales.stream().mapToDouble(Sale::revenue).sum();
        System.out.printf("Total revenue: %,.0f%n", total);

        Map<String, Double> byRegion = sales.stream()
                .collect(Collectors.groupingBy(Sale::region, Collectors.summingDouble(Sale::revenue)));
        byRegion.entrySet().stream()
                .sorted(Map.Entry.<String, Double>comparingByValue().reversed())
                .forEach(e -> System.out.printf("%-6s %,12.0f%n", e.getKey(), e.getValue()));

        String bestProduct = sales.stream()
                .collect(Collectors.groupingBy(Sale::product, Collectors.summingInt(Sale::units)))
                .entrySet().stream()
                .max(Map.Entry.comparingByValue())
                .map(Map.Entry::getKey)
                .orElse("none");
        System.out.println("Most units sold: " + bestProduct);

        List<String> bigDeals = sales.stream()
                .filter(s -> s.revenue() >= 200_000)
                .sorted(Comparator.comparingDouble(Sale::revenue).reversed())
                .map(s -> s.region() + "/" + s.product())
                .toList();
        System.out.println("Big deals: " + bigDeals);
    }
}
Output
Total revenue: 1,298,000
South       535,000
North       491,000
West        272,000
Most units sold: Phone
Big deals: [South/Phone, North/Laptop, North/Phone]

Other ways to create streams#

Java
Stream<String> s1 = Stream.of("a", "b", "c");
IntStream s2 = Arrays.stream(new int[]{1, 2, 3});
Stream<Integer> evens = Stream.iterate(0, n -> n + 2).limit(5);        // 0 2 4 6 8
Stream<Integer> upTo = Stream.iterate(1, n -> n <= 100, n -> n * 3);  // 1 3 9 27 81
Stream<Double> randoms = Stream.generate(Math::random).limit(3);
try (Stream<String> lines = Files.lines(Path.of("data.txt"))) { ... }  // close file streams!

iterate and generate can produce infinite streams; always bound them with limit or takeWhile.

Parallel streams#

list.parallelStream() (or .parallel()) splits work across CPU cores using the common fork/join pool. It can speed up CPU-heavy work on large data, but:

  • For small collections or cheap operations it is often slower.
  • Lambdas must be stateless and side-effect-free; mutating shared collections from a parallel stream causes data races.
  • Ordered operations (findFirst, sorted, forEachOrdered) reduce the benefit.

Measure before using it.

Best practices#

  • Keep lambdas in pipelines short and side-effect-free. Don't add to external lists inside forEach or map; use collect/toList.
  • One operation per line makes pipelines easy to read and debug.
  • Prefer method references (Employee::name) where they read well.
  • A plain for loop is fine, and sometimes clearer, especially with early exits, checked exceptions or index-based logic.
  • Use primitive streams (mapToInt) for numeric work.

Common mistakes#

  • Forgetting the terminal operation, so nothing happens.
  • Reusing a consumed stream (IllegalStateException).
  • Calling .get() on an Optional from findFirst/max without checking.
  • Collectors.toMap with duplicate keys and no merge function.
  • Trying to modify the source collection inside the pipeline.
  • Expecting toList() results to be mutable.

What's next#

Several stream operations returned Optional. Next we look at Optional, Java's tool for representing "maybe there's a value", without null surprises.

Check your understanding

Quick quiz

0/3 answered
  1. 1.What does this print? Stream.of(1, 2, 3).filter(n -> { System.out.print(n); return n > 1; });

  2. 2.Which collector groups employees into a Map<String, List<Employee>> by department?

  3. 3.What happens if you call a terminal operation on a stream twice?

Finished reading?

Mark this lesson complete to track your progress.