The Streams API
Build pipelines with filter, map, sorted and reduce, then collect with groupingBy, joining and more.
The Streams API (java.util.stream, Java 8+) lets you process collections declaratively: say what you want ("the names of active users over 18, sorted"), not how to loop. Combined with lambdas, streams make data processing shorter, clearer and easy to parallelise.
Loop vs stream#
Anatomy of a stream pipeline#
- Source: a collection (
list.stream()), an array (Arrays.stream(arr)),Stream.of(...), a range (IntStream.range), a file (Files.lines)... - Intermediate operations return a new stream:
filter,map,sorted,distinct,limit... They are lazy: nothing happens yet. - Terminal operation produces a result or side effect:
toList,collect,forEach,count,reduce,findFirst... This triggers the processing.
Important properties:
- A stream doesn't modify its source.
- A stream can be consumed only once.
- Elements flow through the pipeline one at a time, so with
limitorfindFirst, much of the source may never be processed.
Elements 5 and 6 are never touched: once limit(2) is satisfied, the pipeline stops.
Core intermediate operations#
Terminal operations#
findFirst, max and min return an Optional, because the stream might be empty (next lesson). reduce(identity, accumulator) combines all elements into one value.
Primitive streams: IntStream, LongStream, DoubleStream#
For numbers, primitive streams avoid boxing and add handy maths methods:
Collectors: building results#
collect(...) with the Collectors utility class builds lists, sets, maps, strings and statistics:
stream.toList()(Java 16+) returns an unmodifiable list and is the most concise option.collect(Collectors.toList())returns a mutableArrayListin practice, but that is not guaranteed.
toMap with duplicate keys throws IllegalStateException. Supply a merge function to resolve them, e.g. Collectors.toMap(Employee::dept, Employee::salary, Double::sum).
A worked example: sales report#
Other ways to create streams#
iterate and generate can produce infinite streams; always bound them with limit or takeWhile.
Parallel streams#
list.parallelStream() (or .parallel()) splits work across CPU cores using the common fork/join pool. It can speed up CPU-heavy work on large data, but:
- For small collections or cheap operations it is often slower.
- Lambdas must be stateless and side-effect-free; mutating shared collections from a parallel stream causes data races.
- Ordered operations (
findFirst,sorted,forEachOrdered) reduce the benefit.
Measure before using it.
Best practices#
- Keep lambdas in pipelines short and side-effect-free. Don't add to external lists inside
forEachormap; usecollect/toList. - One operation per line makes pipelines easy to read and debug.
- Prefer method references (
Employee::name) where they read well. - A plain
forloop is fine, and sometimes clearer, especially with early exits, checked exceptions or index-based logic. - Use primitive streams (
mapToInt) for numeric work.
Common mistakes#
- Forgetting the terminal operation, so nothing happens.
- Reusing a consumed stream (
IllegalStateException). - Calling
.get()on anOptionalfromfindFirst/maxwithout checking. Collectors.toMapwith duplicate keys and no merge function.- Trying to modify the source collection inside the pipeline.
- Expecting
toList()results to be mutable.
What's next#
Several stream operations returned Optional. Next we look at Optional, Java's tool for representing "maybe there's a value", without null surprises.
Check your understanding
Quick quiz
1.What does this print?
Stream.of(1, 2, 3).filter(n -> { System.out.print(n); return n > 1; });2.Which collector groups employees into a
Map<String, List<Employee>>by department?3.What happens if you call a terminal operation on a stream twice?
Finished reading?
Mark this lesson complete to track your progress.