1 of 71

Oilpan’s Performance

& Shipping Plan

haraken@chromium.org

2 of 71

Oilpan team

  • Active members:
  • haraken@
  • sof@ (Opera)
  • tkent@
  • keishi@
  • yutak@
  • peria@
  • Alumni:
  • ager@, erikcorry@, vegorov@, wibling@, zerny@, kouhei@

3 of 71

Current status

  • We have enabled Oilpan by default for all modules/
  • Oilpan’s infrastructure is pretty stabilized

  • We have not yet enabled Oilpan for the Node hierarchy (and a lot of objects tightly coupled with the Node hierarchy)

  • This session shows the latest performance/memory results for the Node hierarchy and a shipping plan

4 of 71

Agenda

  1. Quick overview of Oilpan
  2. Performance & Memory results
  3. Shipping plan
  4. Optimization plan

5 of 71

Quick overview of Oilpan

6 of 71

What is Oilpan?

  • Oilpan is a project to replace reference counting in Blink with a GC

  • If you are not familiar with Oilpan, see this slide and this slide

7 of 71

What is Oilpan?

  • Advantages:
  • Better security (Use-after-frees are gone)
  • Better programmability (Developers don’t need to worry about reference cycles/memory leaks)
  • Better support in developer tools (Developer tools can get a full object graph that crosses over V8 and Blink)

  • If performance is fine, there is no reason we reject Oilpan :)

8 of 71

Oilpan GC

  • Oilpan GC is implemented as a parallel mark & sweep GC
  • Parallel marking is disabled at the moment (for an interesting reason explained later)
  • Parallel sweeping (2 threads) is enabled

9 of 71

Oilpan GC

  • Oilpan GC is a stop-the-world GC
  • All threads (worker threads, database threads etc) need to stop
  • Accurately speaking, a thread doesn’t need to “stop” but just needs to be in a “safe point”
    • “Safe point”: A place where it’s guaranteed that the thread doesn’t allocate anything on the heap
    • e.g., database I/O, file I/O

10 of 71

A relationship between Oilpan GC & V8 GC

11 of 71

A relationship between Oilpan GC & V8 GC

  • V8 GC involves only one thread (a.k.a. Isolate)
  • A V8 heap is completely independent among threads

  • Oilpan GC involves all threads
  • An Oilpan heap is shared among all threads

12 of 71

Precise GC & Conservative GC

  • Oilpan runs a precise GC for on-heap pointers and a conservative GC for on-stack pointers

class A : public GarbageCollected<A> {

void foo() {

B* b = …; // This pointer is found conservatively

}

void trace(Visitor* visitor) { visitor->trace(m_b); }

Member<B> m_b; // This pointer is traced precisely

};

13 of 71

Precise GC & Conservative GC

  • You may worry about the conservativeness… but here is a trick :)
  • The key observation is that Blink’s stack becomes empty when control gets back to the end of an event loop

14 of 71

Precise GC & Conservative GC

  • If we schedule a GC at the end of an event loop, we can run a precise GC without any conservativeness
  • Oilpan tries its best to schedule a precise GC at the end of an event loop
  • Oilpan forces a conservative GC during an event loop only when necessary

15 of 71

Summary

  • Oilpan GC is implemented as a parallel mark & sweep GC
  • Oilpan GC involves all threads in Blink
  • Oilpan tries its best to schedule a precise GC at the end of an event loop
  • Oilpan forces a conservative GC during an event loop only when necessary

16 of 71

Performance & Memory results

17 of 71

Metrics

  • We use the following metrics to evaluate Oilpan GC
    • Execution time
    • Pause time
    • Peak memory usage

18 of 71

Success criteria

  • Execution time
  • Oilpan should be faster than reference counting
  • Pause time
  • Oilpan shouldn’t break 60 FPS
  • Peak memory usage
  • Oilpan shouldn’t increase peak memory usage compared to reference counting

19 of 71

Notes

  • Since Oilpan is a fundamental change to the Blink infrastructure, the result is a mixed bag:
  • Some benchmarks gets better
  • Some benchmarks gets worse

  • The following results are average numbers of given benchmarks
  • You can see the full results as of 2014 October in this spreadsheet

20 of 71

Metrics

  • Execution time
  • Pause time
  • Peak memory usage

21 of 71

Execution time (Blink perf)

  • Blink perf is a set of micro benchmarks for DOM

  • Results:
  • Nexus7: 0.6% better
  • Mac: 0.2% better
  • Linux: 4.8% worse

22 of 71

Execution time (Page cycler)

  • Page cycler is a set of real-world websites to measure page loading time

  • Results:
  • Nexus7: 6.0% better
  • Mac: 5.1% worse
  • Linux: 3.4% worse

23 of 71

Execution time (Speedometer)

  • Speedometer is a set of real-world benchmarks to measure responsiveness of user interactions like editing

  • Results:
  • Nexus7: 2.2% better
  • Mac: 0.6% better
  • Linux: 0.0%

24 of 71

Execution time (Summary)

  • In Nexus7, Oilpan performs slightly better than non-Oilpan

  • In Mac, Oilpan performs as well as non-Oilpan

  • In Linux, Oilpan performs slightly worse than non-Oilpan
  • I will explain the reason later

25 of 71

Metrics

  • Execution time
  • Pause time
  • Peak memory usage

26 of 71

Pause time (Smoothness)

  • Smoothness is a set of real-world benchmarks to measure rendering jankiness

  • Results for smoothness benchmarks except animation benchmarks:
  • Nexus7: 0.3% worse
  • Mac: N/A
  • Linux: 0.7% better

27 of 71

Pause time (Smoothness)

  • However, we observe a significant regression in smoothness.tough_animation_cases

<= Distribution of vsync numbers

28 of 71

Pause time (Web crawling)

  • I manually crawled Facebook, Twitter and Google services and measured the distribution of pause times of Oilpan GC

  • Our goal is to make each Oilpan GC finish in 10 ms

29 of 71

Pause time (Web crawling, Linux)

Pause time

Oilpan precise GC

Oilpan conservative GC

V8 minor GC

V8 major GC

0 - 10 ms

75%

25%

79%

18%

10 - 20 ms

21%

75%

19%

40%

20 - 30 ms

4%

0%

2%

9%

30 - 40 ms

0%

0%

0%

10%

40 - 50 ms

0%

0%

0%

6%

50 - 100 ms

0%

0%

0%

7%

100 ms -

0%

0%

0%

10%

Total occurrences

264

4

274

72

30 of 71

Pause time (Web crawling, Nexus7)

Pause time

Oilpan precise GC

Oilpan conservative GC

V8 minor GC

V8 major GC

0 - 10 ms

20%

0%

29%

0%

10 - 20 ms

23%

0%

35%

0%

20 - 30 ms

7%

0%

17%

3%

30 - 40 ms

6%

0%

5%

8%

40 - 50 ms

7%

0%

6%

2%

50 - 100 ms

23%

0%

8%

25%

100 ms -

14%

0%

0%

62%

Total occurrences

121

0

127

40

31 of 71

Pause time (Summary)

  • A conservative GC is rarely triggered

  • In Linux, 96% of Oilpan GCs finish in 20 ms
  • Looks OK but wants to improve more

  • In Nexus7, only 43% of Oilpan GCs finish in 20 ms and 37% of Oilpan GCs take more than 50 ms
  • Looks unacceptable

32 of 71

Pause time (Summary)

  • That being said, Oilpan GC looks not as bad as V8 GC
  • The pause time distribution of Oilpan GC is simliar to the pause time distribution of V8 minor GC
  • The pause time distribution of Oilpan GC is much better than the pause time distribution of V8 major GC

  • Either way, the pause time is a problem

33 of 71

Metrics

  • Execution time
  • Pause time
  • Peak memory usage

34 of 71

Peak memory usage

  • I measured peak memory usage in memory.tough_dom_memory_cases, memory.typical_mobile_sites and Page cycler

  • The peak memory usage is almost the same between Oilpan and non-Oilpan in all of Nexus7, Mac and Linux

35 of 71

Peak memory usage (Summary)

  • This result looks good and makes sense:
  • According to hajimehoshi’s investigation, Blink consumes only 5% of the total memory of the renderer process
    • V8 and GPU are the main memory consumer
  • Even if Oilpan increases memory usage by 20%, it will just increase the total memory usage by 1%

36 of 71

Summary

  • Execution time is not a problem
  • Except for some specific benchmarks and Linux

  • Pause time is a problem
  • Especially for animation benchmarks

  • Peak memory usage is not a problem

37 of 71

Shipping plan

38 of 71

Current status

  • In order to ship Oilpan for the Node hierarchy, we need to:
  • address performance regressions observed in specific benchmarks
  • reduce pause time

39 of 71

Current status

  • Overall we’re getting close!
  • We will be able to ship Oilpan for the Node hierarchy in a couple of months
  • At the very least, it won’t happen that we give up shipping and revert the Node hierarchy back to a reference counting world

40 of 71

Shipping plan

  • So we’re close to ship it, but the challenging part is how to ship it

  • The Node hierarchy is huge
  • Node, CSS, SVG, RenderObject, AXObjects, Frame, Window, Page, ExecutionContext, ...
  • We cannot simply drop WillBe types in one go

41 of 71

Shipping plan

42 of 71

Summary

  • In short term, we are planning to:
  • start shipping Oilpan for core/ objects independent from the Node hierarchy
  • fix performance issues of the Node hierarchy

  • We are planning to give the first shot of enabling Oilpan for the Node hierarchy in 2015 Q1

  • In any case, we won’t do anything that regresses performance

43 of 71

Optimization plan

44 of 71

What do we need to optimize?

  • In order to ship Oilpan for the Node hierarchy, we need to:
  • address performance regressions observed in specific benchmarks
  • reduce pause time

  • In other words, we need to optimize:
  • Execution time
  • Pause time

45 of 71

What do we need to optimize?

  • Execution time
  • Pause time

46 of 71

About execution time

  • It is important to understand:
  • Factors that make Oilpan better than reference counting
  • Factors that make Oilpan worse than reference counting

47 of 71

Factors that make Oilpan better

  • Oilpan eliminates the overhead of incrementing/decrementing reference counters
  • In particular, Oilpan eliminates all on-stack RefPtrs

void replaceChild(PassRefPtrWillBeRawPtr<Node> child, ...) {

RefPtrWillBeRawPtr<Node> protect(this); // This will be removed

RefPtrWillBeRawPtr<Node> next = child->nextSibling();

…;

}

48 of 71

Factors that make Oilpan better

  • Oilpan eliminates destructors
  • To achieve this goal, it is important to move more things to Oilpan

class Foo : public RefCountedWillBeGarbageCollected() {

#if !ENABLE(OILPAN)

~Foo();

#endif

RefPtrWillBeMember<Bar> m_bar; // RefPtr needs a destructor

};

49 of 71

Factors that make Oilpan worse

  • Oilpan introduces the overhead of marking & sweeping objects

  • Oilpan makes memory locality worse

50 of 71

Things we’ve often observed

  • We’ve often observed that a benchmark regresses by 20% even though a GC spends only 5% of the total time
  • In most cases, a GC itself is not a main bottleneck

  • A problem exists not inside a GC but outside the GC

  • Memory locality is one of the “outside” reasons

51 of 71

Memory locality is important

  • In reference counting:
  • Objects that become unreachable are freed promptly
  • Most objects are short-lived
  • Thus an allocator can reuse memory efficiently for those short-lived objects

  • In a GC:
  • Objects are not freed until a GC is triggered
  • An allocator cannot reuse memory efficiently

52 of 71

Example

// JS benchmark

for (var i = 0; i < 100000; i++) { div.innerHTML = “foo”; }

// Blink code

void Element::setInnerHTML(String str) {

RefPtrWillBeRawPtr<DocumentParser> parser = DocumentParser::create(str);

…;

}

  • In reference counting, all DocumentParser objects are (expected to be) allocated at the same address
  • In a GC, DocumentParser objects created between two GCs are allocated at different addresses

53 of 71

Memory locality is important

  • This problem is inevitable for a GC (to some extent)
  • This often happens in micro benchmarks
  • This also happens in real-world benchmarks (and it’s hard to identify the exact objects causing the problem)
    • e.g., HTMLToken, InterpolableValue, etc

  • Oilpan implements various optimizations to improve the memory locality as much as possible

54 of 71

Memory locality is important

  • Oilpan provides multiple heaps:
  • Dedicated heaps for frequently used types of objects (e.g., Node, CSSValue, RenderObject)
  • A dedicated heap for collections (e.g., Vector, HashTable)
    • This heap supports prompt destruction to support resizing of collections
  • Per-object-size heaps
    • 1 - 3 words, 4 - 7 words, 8 - 15 words, more words

55 of 71

Memory locality is important

  • FIXME?: Oilpan disables parallel marking for now
  • Doing parallel marking improves performance of a GC itself but regresses performance of mutator execution
  • This is because marking threads set marks bits on objects and pollute CPU caches

  • We’re planning to implement bitmap marking to prevent the marking threads from touching the objects

56 of 71

Memory locality is important

  • Remember the performance result for Blink perf:
  • Nexus7: 0.6% better
  • Mac: 0.2% better
  • Linux: 4.8% worse

  • In general, results in Linux are worse than results in Mac and Nexus7
  • Why?

57 of 71

Memory locality is important

  • This can be explained by memory locality
  • Nexus7 is not sensitive to memory locality because the hardware is not as smart as desktop machines
  • Mac is not sensitive to memory locality because the kernel is not as smart as Linux

  • BTW, this is a good news for us because we’re focusing on mobile :)

58 of 71

Plans to improve memory locality

  • Move more things to Oilpan and control the locality

  • It is crazy to mix three allocators in Blink; it makes it harder to control the locality
  • tc-malloc
  • PartitionAlloc
  • Oilpan

59 of 71

Plans to improve memory locality

  • Moving more things to Oilpan will also help speed up the sweeping phase because it eliminates more destructors

class Foo : public GarbageCollectedFinalized<Foo> {

~Foo();

String m_str; // StringImpl is refcounted and needs a destructor

};

  • If we move Strings to Oilpan, we can remove the destructor

60 of 71

Plans to improve memory locality

  • Introduce a generational GC?
  • Probably yes, but no at least in short term
  • There are other low-hanging fruits we should work on first
  • A generational GC is not always a “silver bullet”
    • It adds extra overhead (e.g., write barriers)
    • c.f., JavaScriptCore experimented a generational GC and then decided to go with a parallel mark & sweep GC

61 of 71

What do we need to optimize?

  • Execution time
  • Pause time

62 of 71

Plans to improve pause time

  • Incremental sweeping
  • A mutator can resume its execution immediately after the marking phase
  • A pause time is reduced to a marking time (this will reduce the pause time by more than 50%)

63 of 71

Plans to improve pause time

  • Integrate Oilpan GC into the Blink scheduler
  • Most Oilpan GCs are scheduled at the end of an event loop
  • The Blink scheduler can determine the timing of Oilpan GCs

64 of 71

Plans to improve pause time

  • Oilpan GC is friendly with the Blink scheduler
  • It is OK to not schedule Oilpan GC immediately
  • For example, it is OK to schedule user-input tasks first, if the tasks are not likely to allocate many objects
  • Even if the expectation is violated, nothing serious will happen; Oilpan just triggers a conservative GC

  • By scheduling Oilpan GC wisely, we can reduce the impact of the GC pauses

65 of 71

Plans to improve pause time

  • In reality, animation benchmarks are the only regressing benchmarks in Nexus7

  • We can (partially or entirely) drop Oilpan from core/animations/ (at least in short term)
  • core/animations/ can keep using RefPtrs
  • c.f., We decided not to introduce Oilpan for realtime WebAudio threads

66 of 71

Summary

  • To optimize execution time:
  • Improve memory locality
  • Identify performance issues of specific regressing benchmarks
  • To optimize pause time:
  • Implement incremental sweeping
  • Integrate Oilpan GC into the Blink scheduler
  • Drop Oilpan from core/animations/
  • Optimize GC parameters for mobile

67 of 71

Summary

  • I think these optimizations will address most of the performance issues blocking the Node hierarchy

  • I don’t think these optimizations take a couple of months

  • I think we can give the first shot of shipping Oilpan for the Node hierarchy in 2015 Q1

68 of 71

Conclusions

69 of 71

Conclusions

  • Execution time is not a problem
  • Except for some specific benchmarks and Linux

  • Pause time is a problem
  • Especially for animation benchmarks

  • Peak memory usage is not a problem

70 of 71

Conclusions

  • It took a long time, but we’re getting close!

  • In short term, we are planning to:
  • enable Oilpan by default for core/ objects independent from the Node hierarchy
  • fix performance issues of the Node hierarchy

  • We are planning to give the first shot of shipping Oilpan for the Node hierarchy in 2015 Q1

71 of 71

Q & A