I'm not sure Segment is really helping performance in real-world use cases. This was noticed recently in https://github.com/functional-streams-for-scala/fs2/issues/1160, where even with some optimizations in 1.0.0-SNAPSHOT, a real world use case is about 2x slower under 1.0 than under 0.9.
Last time I did some benchmarking of 0.9 vs 0.10, I found that while 0.10 was significantly faster than 0.9 in most use cases, writing algorithms in terms of chunks instead of segments was generally faster (https://speakerdeck.com/mpilquist/fs2-internals?slide=24). Segment did win with certain chunk patterns though. Those results were with 0.10.0-M8. Running the same tests against 1.0.0-SNAPSHOT should favor the chunk backed algorithms even more.
Algorithms are often much harder to implement in terms of Segment than in terms of Chunk. Users often don't know which to implement and we typically offer the advice that Segment based algorithms will be faster. However the take benchmark shows that this isn't always true. Given the complexity of working with Segment, I feel like the performance gains need to be great in order to outweigh the complexity.
Like I said here, we shouldn't make any decisions without experiments. I'd appreciate feedback though. One option could be supporting both Segment and Chunk internally instead of lifting every chunk in to a segment? Or keep Segment but deprecate it?
I think for 1.0 I would made segment just private, and thus force people to use chunk.
if we were to deprecate Segment, will be need to rewrite every op to support chunking explicitly?
We'd need to rewrite existing operations to use unconsChunk instead of uncons and in doing so we'd lose operation fusion on segments. However I don't suspect we'd care in real use cases -- e.g., even micro benchmarks like s.map(f).map(g).map(h) are likely to be faster with chunks as opposed to segments, as the cost of copying is cheap compared to constant factors with segment (or at least that's the argument, we'd need to confirm).
I also believe that because the chunk will be strict, the interpreter will be easier imho
I'm referring to this comment
Remember in 0.9, we had adhoc special logic for preserving chunkiness -- basically everything had to go through mapChunks in order to preserve chunkiness. Segment solves all that as long as you don't try to observe chunk structure (though use of s.pull.unconsChunk).
I'm aware we'd lose fusion, however it's unclear if we can preserve the property of preserving chunkiness without having to have too much custom logic everywhere. Is unconsChunk enough?
yeah, perhaps we just need to remove segment from API, however it has to stay in implementation .
Hm, could you elaborate some on why we might lose chunkiness? Seems like implementing pulls in terms of unconsChunk guarantees no loss of chunkiness by construction -- just with overhead of copying subchunks sometimes.
One major criticism of the current segment based approach is that you lose chunkiness on almost any operation (but you keep segmentiness) -- e.g., (s: Segment[Int, Unit]).map(f) results in loss of chunkiness b/c each element is emitted individually. My recent changes to Segment#unconsChunk attempt to overcome this but only via a bunch of copying, which means loss of specialized chunk types. E.g., #1164 results in loss of underlying chunk types.
Hm, could you elaborate some on why we might lose chunkiness? Seems like implementing pulls in terms of unconsChunk guarantees no loss of chunkiness by construction -- just with overhead of copying subchunks sometimes.
That comment is from you, so I don't know :P
I'm not opposed to removing Segment. It's hard to explain and some things are harder to implement, but it's worth it for perf, so if we lose perf I'm ok with getting rid of it.
Seems like implementing pulls in terms of unconsChunk guarantees no loss of chunkiness by construction -- just with overhead of copying subchunks sometimes.
Yeah, that would have been my guess, so I'm fine with it :)
@mpilquist any idea how Stream#flatMap will be defined with Chunk only instead of lazy segment?
as for the copying the specialised chunks, is there a ways to have "specialised" def copy on chunk types ?
any idea how Stream#flatMap will be defined with Chunk only instead of lazy segment?
@pchlupacek Basically the same as now -- we'd need a foldRightLazy[B](z: B)(f: (A, => B) => B): B method on Chunk[A]. We could lift the implementation of Segment#foldRightLazy and remove the outer layer of recursion on segments. Stream#flatMap doesn't use Segment#flatMap currently.
as for the copying the specialised chunks, is there a ways to have "specialised" def copy on chunk types ?
Right, we'd use normal subtype polymorphism to provide specialized implementations of various chunk operations.
@mpilquist If that would work I am fine with that.
Could we brainstorm some performance tests we'd need to run in order to make a decision? Some general purpose tests run against streams of various chunk levels, some tests specifically around moving big chunky byte arrays around, etc.? I don't mind experimenting with a branch which reduces importance of Segment but I'd really like a way to test assumptions.
I would love to have something like
val byteSource: Stream[F,Byte] = /** chunked stream of bytes i.e. packets from socket with various sizes **/
val repacketize: PIpe[F, Byte, Byte] = /** rpacketize 1) make smaller chunks, 2) concat chunks, various sizes related to size of !byteSource **/
// then
byteSource through repacketize
that covers what you most often do in http/socket environments
@mpilquist I can’t remember the comment right now (on my phone), but I suggested a case on the ticket you linked where segments should be vastly, vastly slower than chunks. I think it was something like reading json off disk, parsing it with jawn streaming, modifying the json lightly, then marshaling and emitting the results via http4s.
I started some benchmarks here: https://github.com/functional-streams-for-scala/fs2-benchmarks
The first one I added was reading a large json file from disk, parsing to strings, and splitting on newlines. 0.10 and 1.0 are about 2.1x faster than 0.9 at that test. See the jmh-result.json files in the 0.9, 0.10, and 1.0 directories.
thats excellent @mpilquist
Interesting. I wonder what would more accurately capture the archetype of the massive performance hit we saw in quasar. @wemrysi any ideas?
@djspiewak We'd need a high cost-per-chunk consumer. Maybe simulating something that processes a chunk at a time via a relatively expensive operation per chunk. Even a naive sleep would illustrate it, though maybe trivially so.
The cost proportional to chunk size would have to increase at a slower rate than the cost proportional to the number of chunks.
I have a work in progress branch that removes Segment here: https://github.com/functional-streams-for-scala/fs2/compare/series/1.0...mpilquist:topic/back-to-the-chunk
Note a bunch of stuff is left commented out but enough is in place for us to run the benchmarks.
A quick run through the fs2-benchmarks project shows improvements in both test cases:
1.0 with Segment
info] Benchmark Mode Cnt Score Error Units
[info] TextParsingBenchmark.parseAndMapBigFileSync avgt 4 3237.255 ± 2004.354 ms/op
[info] TextParsingBenchmark.parseBigFileSync avgt 4 2830.396 ± 3118.065 ms/op
1.0 without Segment
[info] TextParsingBenchmark.parseAndMapBigFileSync avgt 4 3186.769 ± 1819.525 ms/op
[info] TextParsingBenchmark.parseBigFileSync avgt 4 2197.027 ± 227.529 ms/op
This certainly points towards removing Segment though we likely need some more benchmarks before making a decision.
That branch has been updated -- all tests are passing now without Segment involved at all. One thing that's kind of nice is that Stream.segment and Pull.segment still exist -- they are basically unfolds of the segment in to chunks of a max size. This unfold always existed inside the interpreter (and we never came up with a good way to parameterize the max chunk size and max step size). With this approach, the user can specify those params when calling Stream.segment.
I ran the existing performance tests against both versions and it appears that the chunk based version, even without specialized implementations of new chunk operations, wins in almost every test and when it doesn't win, it's pretty close. If no objections, I'll prepare a PR.
1.0.0-SNAPSHOT with Segment
Revision: 43fc17bcc01a27603a057ac3dd6340b9acb8a0ad
[info] StreamPerformanceSpec:
[info] Stream Performance
[info] left-associated ++
[info] - 2 (427 milliseconds)
[info] - 3 (2 milliseconds)
[info] - 100 (5 milliseconds)
[info] - 200 (9 milliseconds)
[info] - 400 (12 milliseconds)
[info] - 800 (15 milliseconds)
[info] - 1600 (11 milliseconds)
[info] - 3200 (14 milliseconds)
[info] - 6400 (20 milliseconds)
[info] - 12800 (19 milliseconds)
[info] - 25600 (38 milliseconds)
[info] - 51200 (77 milliseconds)
[info] - 102400 (69 milliseconds)
[info] right-associated ++
[info] - 2 (2 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (1 millisecond)
[info] - 200 (1 millisecond)
[info] - 400 (2 milliseconds)
[info] - 800 (2 milliseconds)
[info] - 1600 (3 milliseconds)
[info] - 3200 (3 milliseconds)
[info] - 6400 (6 milliseconds)
[info] - 12800 (7 milliseconds)
[info] - 25600 (15 milliseconds)
[info] - 51200 (54 milliseconds)
[info] - 102400 (56 milliseconds)
[info] left-associated flatMap 1
[info] - 2 (21 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (7 milliseconds)
[info] - 200 (12 milliseconds)
[info] - 400 (12 milliseconds)
[info] - 800 (13 milliseconds)
[info] - 1600 (16 milliseconds)
[info] - 3200 (29 milliseconds)
[info] - 6400 (39 milliseconds)
[info] - 12800 (55 milliseconds)
[info] - 25600 (44 milliseconds)
[info] - 51200 (132 milliseconds)
[info] - 102400 (265 milliseconds)
[info] left-associated eval() ++ flatMap 1
[info] - 2 (13 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (6 milliseconds)
[info] - 200 (4 milliseconds)
[info] - 400 (11 milliseconds)
[info] - 800 (10 milliseconds)
[info] - 1600 (20 milliseconds)
[info] - 3200 (41 milliseconds)
[info] - 6400 (65 milliseconds)
[info] - 12800 (45 milliseconds)
[info] - 25600 (90 milliseconds)
[info] - 51200 (242 milliseconds)
[info] - 102400 (727 milliseconds)
[info] right-associated flatMap 1
[info] - 2 (1 millisecond)
[info] - 3 (1 millisecond)
[info] - 100 (1 millisecond)
[info] - 200 (2 milliseconds)
[info] - 400 (2 milliseconds)
[info] - 800 (2 milliseconds)
[info] - 1600 (3 milliseconds)
[info] - 3200 (8 milliseconds)
[info] - 6400 (11 milliseconds)
[info] - 12800 (22 milliseconds)
[info] - 25600 (41 milliseconds)
[info] - 51200 (120 milliseconds)
[info] - 102400 (155 milliseconds)
[info] right-associated eval() ++ flatMap 1
[info] - 2 (3 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (2 milliseconds)
[info] - 200 (2 milliseconds)
[info] - 400 (3 milliseconds)
[info] - 800 (6 milliseconds)
[info] - 1600 (7 milliseconds)
[info] - 3200 (13 milliseconds)
[info] - 6400 (22 milliseconds)
[info] - 12800 (43 milliseconds)
[info] - 25600 (124 milliseconds)
[info] - 51200 (488 milliseconds)
[info] - 102400 (324 milliseconds)
[info] left-associated flatMap 2
[info] - 2 (2 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (1 millisecond)
[info] - 200 (2 milliseconds)
[info] - 400 (2 milliseconds)
[info] - 800 (4 milliseconds)
[info] - 1600 (6 milliseconds)
[info] - 3200 (16 milliseconds)
[info] - 6400 (28 milliseconds)
[info] - 12800 (92 milliseconds)
[info] - 25600 (114 milliseconds)
[info] - 51200 (230 milliseconds)
[info] - 102400 (775 milliseconds)
[info] right-associated flatMap 2
[info] - 2 (3 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (1 millisecond)
[info] - 200 (1 millisecond)
[info] - 400 (2 milliseconds)
[info] - 800 (2 milliseconds)
[info] - 1600 (2 milliseconds)
[info] - 3200 (7 milliseconds)
[info] - 6400 (10 milliseconds)
[info] - 12800 (20 milliseconds)
[info] - 25600 (41 milliseconds)
[info] - 51200 (199 milliseconds)
[info] - 102400 (155 milliseconds)
[info] transduce (id)
[info] - 2 (31 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (3 milliseconds)
[info] - 200 (4 milliseconds)
[info] - 400 (5 milliseconds)
[info] - 800 (10 milliseconds)
[info] - 1600 (9 milliseconds)
[info] - 3200 (9 milliseconds)
[info] - 6400 (17 milliseconds)
[info] - 12800 (26 milliseconds)
[info] - 25600 (31 milliseconds)
[info] - 51200 (298 milliseconds)
[info] - 102400 (110 milliseconds)
[info] bracket + handleErrorWith (1)
[info] - 2 (19 milliseconds)
[info] - 3 (2 milliseconds)
[info] - 100 (28 milliseconds)
[info] - 200 (28 milliseconds)
[info] - 400 (42 milliseconds)
[info] - 800 (28 milliseconds)
[info] - 1600 (33 milliseconds)
[info] - 3200 (51 milliseconds)
[info] - 6400 (57 milliseconds)
[info] - 12800 (88 milliseconds)
[info] - 25600 (175 milliseconds)
[info] - 51200 (320 milliseconds)
[info] - 102400 (764 milliseconds)
[info] chunky flatMap
[info] - 2 (1 millisecond)
[info] - 3 (0 milliseconds)
[info] - 100 (0 milliseconds)
[info] - 200 (1 millisecond)
[info] - 400 (1 millisecond)
[info] - 800 (1 millisecond)
[info] - 1600 (1 millisecond)
[info] - 3200 (2 milliseconds)
[info] - 6400 (3 milliseconds)
[info] - 12800 (8 milliseconds)
[info] - 25600 (12 milliseconds)
[info] - 51200 (21 milliseconds)
[info] - 102400 (43 milliseconds)
[info] chunky map with unconsChunk
[info] - 2 (6 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (1 millisecond)
[info] - 200 (2 milliseconds)
[info] - 400 (2 milliseconds)
[info] - 800 (1 millisecond)
[info] - 1600 (1 millisecond)
[info] - 3200 (2 milliseconds)
[info] - 6400 (4 milliseconds)
[info] - 12800 (5 milliseconds)
[info] - 25600 (8 milliseconds)
[info] - 51200 (8 milliseconds)
[info] - 102400 (8 milliseconds)
[info] ScalaTest
[info] Run completed in 10 seconds, 490 milliseconds.
[info] Total number of tests run: 156
[info] Suites: completed 1, aborted 0
[info] Tests: succeeded 156, failed 0, canceled 0, ignored 0, pending 0
[info] All tests passed.
[info] Passed: Total 156, Failed 0, Errors 0, Passed 156
[success] Total time: 55 s, completed Jul 23, 2018 9:04:18 AM
1.0.0-SNAPSHOT without Segment
Revision: 07d7dd940ff8c9d9ea1d942e7d157c99c0b754ab
[info] Benchmark Mode Cnt Score Error Units
[info] StreamBenchmark.eval_1 thrpt 25 294034.618 ± 12115.299 ops/s
[info] StreamBenchmark.eval_10 thrpt 25 61851.975 ± 6129.738 ops/s
[info] StreamBenchmark.eval_100 thrpt 25 8035.465 ± 158.476 ops/s
[info] StreamBenchmark.eval_1000 thrpt 25 818.947 ± 22.715 ops/s
[info] StreamBenchmark.eval_10000 thrpt 25 77.005 ± 4.671 ops/s
[info] StreamBenchmark.eval_100000 thrpt 25 6.804 ± 0.377 ops/s
[info] StreamBenchmark.leftAssocConcat_10 thrpt 25 284995.409 ± 2689.927 ops/s
[info] StreamBenchmark.leftAssocConcat_100 thrpt 25 33372.259 ± 184.023 ops/s
[info] StreamBenchmark.leftAssocConcat_1000 thrpt 25 3384.512 ± 31.530 ops/s
[info] StreamBenchmark.leftAssocConcat_10000 thrpt 25 306.371 ± 5.404 ops/s
[info] StreamBenchmark.leftAssocConcat_100000 thrpt 25 8.924 ± 0.184 ops/s
[info] StreamBenchmark.leftAssocFlatMap_1 thrpt 25 879159.422 ± 13673.559 ops/s
[info] StreamBenchmark.leftAssocFlatMap_10 thrpt 25 109929.617 ± 1265.226 ops/s
[info] StreamBenchmark.leftAssocFlatMap_100 thrpt 25 10419.919 ± 1345.402 ops/s
[info] StreamBenchmark.leftAssocFlatMap_1000 thrpt 25 932.568 ± 88.116 ops/s
[info] StreamBenchmark.leftAssocFlatMap_10000 thrpt 25 61.024 ± 3.835 ops/s
[info] StreamBenchmark.leftAssocFlatMap_100000 thrpt 25 1.991 ± 0.085 ops/s
[info] StreamBenchmark.rightAssocConcat_1 thrpt 25 848889.241 ± 16994.998 ops/s
[info] StreamBenchmark.rightAssocConcat_10 thrpt 25 266178.821 ± 2384.894 ops/s
[info] StreamBenchmark.rightAssocConcat_100 thrpt 25 32828.111 ± 277.639 ops/s
[info] StreamBenchmark.rightAssocConcat_1000 thrpt 25 3377.183 ± 36.569 ops/s
[info] StreamBenchmark.rightAssocConcat_10000 thrpt 25 306.760 ± 7.912 ops/s
[info] StreamBenchmark.rightAssocConcat_100000 thrpt 25 9.016 ± 0.199 ops/s
[info] StreamBenchmark.rightAssocFlatMap_1 thrpt 25 847688.143 ± 9263.712 ops/s
[info] StreamBenchmark.rightAssocFlatMap_10 thrpt 25 95909.236 ± 8356.185 ops/s
[info] StreamBenchmark.rightAssocFlatMap_100 thrpt 25 9335.221 ± 891.394 ops/s
[info] StreamBenchmark.rightAssocFlatMap_1000 thrpt 25 938.948 ± 78.192 ops/s
[info] StreamBenchmark.rightAssocFlatMap_10000 thrpt 25 67.862 ± 4.415 ops/s
[info] StreamBenchmark.rightAssocFlatMap_100000 thrpt 25 3.038 ± 0.174 ops/s
[info] StreamBenchmark.toVector_1 thrpt 25 863473.290 ± 16674.705 ops/s
[info] StreamBenchmark.toVector_10 thrpt 25 775762.492 ± 84582.212 ops/s
[info] StreamBenchmark.toVector_100 thrpt 25 450680.900 ± 49948.313 ops/s
[info] StreamBenchmark.toVector_1000 thrpt 25 88835.737 ± 1080.181 ops/s
[info] StreamBenchmark.toVector_10000 thrpt 25 10014.986 ± 88.468 ops/s
[info] StreamBenchmark.toVector_100000 thrpt 25 1016.251 ± 4.071 ops/s
[info] StreamBenchmark.unconsPull_256 thrpt 25 11945.749 ± 89.344 ops/s
[info] StreamBenchmark.unconsPull_8 thrpt 25 446.202 ± 1.957 ops/s
[info] StreamBenchmark.emitsThenFlatMap_1 avgt 25 2552.005 ± 20.146 ns/op
[info] StreamBenchmark.emitsThenFlatMap_10 avgt 25 6592.788 ± 71.177 ns/op
[info] StreamBenchmark.emitsThenFlatMap_100 avgt 25 37081.370 ± 6404.854 ns/op
[info] StreamBenchmark.emitsThenFlatMap_1000 avgt 25 315960.990 ± 2119.083 ns/op
[info] StreamBenchmark.emitsThenFlatMap_10000 avgt 25 3438304.943 ± 27530.708 ns/op
[success] Total time: 2221 s, completed Jul 23, 2018 12:43:27 PM
[info] StreamPerformanceSpec:
[info] Stream Performance
[info] left-associated ++
[info] - 2 (427 milliseconds)
[info] - 3 (3 milliseconds)
[info] - 100 (5 milliseconds)
[info] - 200 (9 milliseconds)
[info] - 400 (11 milliseconds)
[info] - 800 (10 milliseconds)
[info] - 1600 (15 milliseconds)
[info] - 3200 (12 milliseconds)
[info] - 6400 (16 milliseconds)
[info] - 12800 (18 milliseconds)
[info] - 25600 (27 milliseconds)
[info] - 51200 (48 milliseconds)
[info] - 102400 (85 milliseconds)
[info] right-associated ++
[info] - 2 (2 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (1 millisecond)
[info] - 200 (1 millisecond)
[info] - 400 (1 millisecond)
[info] - 800 (3 milliseconds)
[info] - 1600 (2 milliseconds)
[info] - 3200 (3 milliseconds)
[info] - 6400 (4 milliseconds)
[info] - 12800 (6 milliseconds)
[info] - 25600 (12 milliseconds)
[info] - 51200 (50 milliseconds)
[info] - 102400 (23 milliseconds)
[info] left-associated flatMap 1
[info] - 2 (16 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (3 milliseconds)
[info] - 200 (6 milliseconds)
[info] - 400 (6 milliseconds)
[info] - 800 (7 milliseconds)
[info] - 1600 (6 milliseconds)
[info] - 3200 (9 milliseconds)
[info] - 6400 (19 milliseconds)
[info] - 12800 (55 milliseconds)
[info] - 25600 (24 milliseconds)
[info] - 51200 (66 milliseconds)
[info] - 102400 (329 milliseconds)
[info] left-associated eval() ++ flatMap 1
[info] - 2 (6 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (3 milliseconds)
[info] - 200 (4 milliseconds)
[info] - 400 (7 milliseconds)
[info] - 800 (7 milliseconds)
[info] - 1600 (5 milliseconds)
[info] - 3200 (8 milliseconds)
[info] - 6400 (14 milliseconds)
[info] - 12800 (53 milliseconds)
[info] - 25600 (46 milliseconds)
[info] - 51200 (197 milliseconds)
[info] - 102400 (298 milliseconds)
[info] right-associated flatMap 1
[info] - 2 (1 millisecond)
[info] - 3 (0 milliseconds)
[info] - 100 (1 millisecond)
[info] - 200 (1 millisecond)
[info] - 400 (2 milliseconds)
[info] - 800 (1 millisecond)
[info] - 1600 (2 milliseconds)
[info] - 3200 (4 milliseconds)
[info] - 6400 (59 milliseconds)
[info] - 12800 (15 milliseconds)
[info] - 25600 (23 milliseconds)
[info] - 51200 (42 milliseconds)
[info] - 102400 (86 milliseconds)
[info] right-associated eval() ++ flatMap 1
[info] - 2 (2 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (2 milliseconds)
[info] - 200 (3 milliseconds)
[info] - 400 (2 milliseconds)
[info] - 800 (2 milliseconds)
[info] - 1600 (5 milliseconds)
[info] - 3200 (10 milliseconds)
[info] - 6400 (38 milliseconds)
[info] - 12800 (80 milliseconds)
[info] - 25600 (189 milliseconds)
[info] - 51200 (84 milliseconds)
[info] - 102400 (261 milliseconds)
[info] left-associated flatMap 2
[info] - 2 (4 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (2 milliseconds)
[info] - 200 (2 milliseconds)
[info] - 400 (3 milliseconds)
[info] - 800 (4 milliseconds)
[info] - 1600 (5 milliseconds)
[info] - 3200 (8 milliseconds)
[info] - 6400 (12 milliseconds)
[info] - 12800 (24 milliseconds)
[info] - 25600 (49 milliseconds)
[info] - 51200 (106 milliseconds)
[info] - 102400 (475 milliseconds)
[info] right-associated flatMap 2
[info] - 2 (2 milliseconds)
[info] - 3 (0 milliseconds)
[info] - 100 (1 millisecond)
[info] - 200 (1 millisecond)
[info] - 400 (2 milliseconds)
[info] - 800 (3 milliseconds)
[info] - 1600 (6 milliseconds)
[info] - 3200 (12 milliseconds)
[info] - 6400 (17 milliseconds)
[info] - 12800 (158 milliseconds)
[info] - 25600 (22 milliseconds)
[info] - 51200 (41 milliseconds)
[info] - 102400 (81 milliseconds)
[info] transduce (id)
[info] - 2 (24 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (4 milliseconds)
[info] - 200 (5 milliseconds)
[info] - 400 (6 milliseconds)
[info] - 800 (10 milliseconds)
[info] - 1600 (9 milliseconds)
[info] - 3200 (9 milliseconds)
[info] - 6400 (16 milliseconds)
[info] - 12800 (24 milliseconds)
[info] - 25600 (29 milliseconds)
[info] - 51200 (56 milliseconds)
[info] - 102400 (121 milliseconds)
[info] bracket + handleErrorWith (1)
[info] - 2 (27 milliseconds)
[info] - 3 (3 milliseconds)
[info] - 100 (28 milliseconds)
[info] - 200 (36 milliseconds)
[info] - 400 (49 milliseconds)
[info] - 800 (30 milliseconds)
[info] - 1600 (34 milliseconds)
[info] - 3200 (44 milliseconds)
[info] - 6400 (46 milliseconds)
[info] - 12800 (65 milliseconds)
[info] - 25600 (122 milliseconds)
[info] - 51200 (361 milliseconds)
[info] - 102400 (577 milliseconds)
[info] chunky flatMap
[info] - 2 (2 milliseconds)
[info] - 3 (1 millisecond)
[info] - 100 (1 millisecond)
[info] - 200 (1 millisecond)
[info] - 400 (1 millisecond)
[info] - 800 (2 milliseconds)
[info] - 1600 (2 milliseconds)
[info] - 3200 (2 milliseconds)
[info] - 6400 (3 milliseconds)
[info] - 12800 (5 milliseconds)
[info] - 25600 (7 milliseconds)
[info] - 51200 (13 milliseconds)
[info] - 102400 (25 milliseconds)
[info] chunky map with uncons
[info] - 2 (11 milliseconds)
[info] - 3 (0 milliseconds)
[info] - 100 (1 millisecond)
[info] - 200 (1 millisecond)
[info] - 400 (1 millisecond)
[info] - 800 (1 millisecond)
[info] - 1600 (1 millisecond)
[info] - 3200 (1 millisecond)
[info] - 6400 (2 milliseconds)
[info] - 12800 (2 milliseconds)
[info] - 25600 (3 milliseconds)
[info] - 51200 (3 milliseconds)
[info] - 102400 (4 milliseconds)
[info] Benchmark Mode Cnt Score Error Units
[info] StreamBenchmark.eval_1 thrpt 25 374270.636 ± 31228.359 ops/s
[info] StreamBenchmark.eval_10 thrpt 25 116815.054 ± 7026.506 ops/s
[info] StreamBenchmark.eval_100 thrpt 25 15149.811 ± 89.715 ops/s
[info] StreamBenchmark.eval_1000 thrpt 25 1569.698 ± 24.531 ops/s
[info] StreamBenchmark.eval_10000 thrpt 25 153.204 ± 1.821 ops/s
[info] StreamBenchmark.eval_100000 thrpt 25 15.702 ± 0.091 ops/s
[info] StreamBenchmark.leftAssocConcat_10 thrpt 25 406724.950 ± 5199.734 ops/s
[info] StreamBenchmark.leftAssocConcat_100 thrpt 25 53267.576 ± 564.152 ops/s
[info] StreamBenchmark.leftAssocConcat_1000 thrpt 25 5298.008 ± 42.260 ops/s
[info] StreamBenchmark.leftAssocConcat_10000 thrpt 25 529.310 ± 3.672 ops/s
[info] StreamBenchmark.leftAssocConcat_100000 thrpt 25 12.527 ± 0.879 ops/s
[info] StreamBenchmark.leftAssocFlatMap_1 thrpt 25 993513.435 ± 10171.055 ops/s
[info] StreamBenchmark.leftAssocFlatMap_10 thrpt 25 162762.607 ± 1157.101 ops/s
[info] StreamBenchmark.leftAssocFlatMap_100 thrpt 25 16446.874 ± 1677.096 ops/s
[info] StreamBenchmark.leftAssocFlatMap_1000 thrpt 25 1745.401 ± 10.547 ops/s
[info] StreamBenchmark.leftAssocFlatMap_10000 thrpt 25 110.389 ± 1.518 ops/s
[info] StreamBenchmark.leftAssocFlatMap_100000 thrpt 25 2.227 ± 0.149 ops/s
[info] StreamBenchmark.rightAssocConcat_1 thrpt 25 979715.011 ± 15720.721 ops/s
[info] StreamBenchmark.rightAssocConcat_10 thrpt 25 376757.010 ± 4181.336 ops/s
[info] StreamBenchmark.rightAssocConcat_100 thrpt 25 53708.221 ± 240.056 ops/s
[info] StreamBenchmark.rightAssocConcat_1000 thrpt 25 5365.081 ± 30.112 ops/s
[info] StreamBenchmark.rightAssocConcat_10000 thrpt 25 533.073 ± 4.313 ops/s
[info] StreamBenchmark.rightAssocConcat_100000 thrpt 25 12.152 ± 1.013 ops/s
[info] StreamBenchmark.rightAssocFlatMap_1 thrpt 25 948534.702 ± 21681.175 ops/s
[info] StreamBenchmark.rightAssocFlatMap_10 thrpt 25 151771.749 ± 1595.922 ops/s
[info] StreamBenchmark.rightAssocFlatMap_100 thrpt 25 15960.979 ± 81.643 ops/s
[info] StreamBenchmark.rightAssocFlatMap_1000 thrpt 25 1630.147 ± 17.754 ops/s
[info] StreamBenchmark.rightAssocFlatMap_10000 thrpt 25 113.195 ± 0.885 ops/s
[info] StreamBenchmark.rightAssocFlatMap_100000 thrpt 25 3.648 ± 0.185 ops/s
[info] StreamBenchmark.toVector_1 thrpt 25 990510.449 ± 5781.730 ops/s
[info] StreamBenchmark.toVector_10 thrpt 25 958756.945 ± 9395.886 ops/s
[info] StreamBenchmark.toVector_100 thrpt 25 671697.597 ± 7995.634 ops/s
[info] StreamBenchmark.toVector_1000 thrpt 25 141227.826 ± 15004.947 ops/s
[info] StreamBenchmark.toVector_10000 thrpt 25 17592.567 ± 307.240 ops/s
[info] StreamBenchmark.toVector_100000 thrpt 25 1794.049 ± 31.311 ops/s
[info] StreamBenchmark.unconsPull_256 thrpt 25 18090.215 ± 127.625 ops/s
[info] StreamBenchmark.unconsPull_8 thrpt 25 2893.764 ± 72.349 ops/s
[info] StreamBenchmark.emitsThenFlatMap_1 avgt 25 1610.681 ± 18.875 ns/op
[info] StreamBenchmark.emitsThenFlatMap_10 avgt 25 3462.339 ± 26.916 ns/op
[info] StreamBenchmark.emitsThenFlatMap_100 avgt 25 20921.488 ± 197.375 ns/op
[info] StreamBenchmark.emitsThenFlatMap_1000 avgt 25 208592.714 ± 36862.482 ns/op
[info] StreamBenchmark.emitsThenFlatMap_10000 avgt 25 2004565.785 ± 28831.666 ns/op
[success] Total time: 2216 s, completed Jul 23, 2018 1:45:14 PM
Sounds good :)
@mpilquist So in theory we should be able to get significantly even more performance under certain cases by implementing the specialized Chunk operations, right?
I think so, though who knows how much. On the branch, I implemented all the new methods directly on Chunk, never specializing in a subtype. Most of them allocate a s.c.m.Buffer builder with a size hint. The primitive aware subtypes should definitely be able to beat this performance wise due to boxing/unboxing. I don't think such optimizations would show up in our benchmarks though, as not many of the stream benchmarks use large primitive chunks.
Most helpful comment
I ran the existing performance tests against both versions and it appears that the chunk based version, even without specialized implementations of new chunk operations, wins in almost every test and when it doesn't win, it's pretty close. If no objections, I'll prepare a PR.