vChewing-macOS

Commit Graph

Author	SHA1	Message	Date
Lukhnos Liu	d064f420e4	Use a parseless phrase db to speed up LM loading We take advantage of the fact that no one is able to modify the phrase databases shipped with the binary (guranteed by macOS's integrity check for notarized apps), and we can simply pre-sort the phrases in the database files. With this change, we can speed up McBopomofo's language model loading during the app initialization by about 500-800x on a 2018 Intel MacBook Pro. The LM loading used to take 300-400 ms, but now it's done within a sub-millisecond range (0.5-0.6 ms). Microbenchmarking shows that ParselessLM is about 16000x faster than FastLM. We amortize the latency during the query time, and even by deferring the parsing, ParselessLM is only ~1.5x slower than FastLM, and both LM classes serve queries unedr 6 microseconds (that's 0.006 ms), which means the tradeoff only contributes to neglible overall latency. This PR requires some small changes to the phrase db cooking scripts. Python 3 is now used and the (value, reading, score) tuples are rearranged to (reading, value, score) and sorted by reading ("key"). A header is added to the phrase databases to call out the fact that these are pre-sorted. clang-format is used to apply WebKit C++ style to the new code. This also applies to KeyValueBlobReader that was added recently. Microbenchmark result below: ``` --------------------------------------------------------------------- Benchmark Time CPU Iterations --------------------------------------------------------------------- BM_ParselessLMOpenClose 17710 ns 17199 ns 33422 BM_FastLMOpenClose 376520248 ns 367526500 ns 2 BM_ParselessLMFindUnigrams 5967 ns 5899 ns 113729 BM_FastLMFindUnigrams 2268 ns 2265 ns 307038 ```	2022-01-15 16:15:02 -08:00

Author

SHA1

Message

Date

Lukhnos Liu

d064f420e4

Use a parseless phrase db to speed up LM loading

We take advantage of the fact that no one is able to modify the phrase
databases shipped with the binary (guranteed by macOS's integrity check
for notarized apps), and we can simply pre-sort the phrases in the
database files.

With this change, we can speed up McBopomofo's language model loading
during the app initialization by about 500-800x on a 2018 Intel MacBook
Pro. The LM loading used to take 300-400 ms, but now it's done within a
sub-millisecond range (0.5-0.6 ms). Microbenchmarking shows that
ParselessLM is about 16000x faster than FastLM. We amortize the latency
during the query time, and even by deferring the parsing, ParselessLM is
only ~1.5x slower than FastLM, and both LM classes serve queries unedr 6
microseconds (that's 0.006 ms), which means the tradeoff only
contributes to neglible overall latency.

This PR requires some small changes to the phrase db cooking scripts.
Python 3 is now used and the (value, reading, score) tuples are
rearranged to (reading, value, score) and sorted by reading ("key"). A
header is added to the phrase databases to call out the fact that these
are pre-sorted.

clang-format is used to apply WebKit C++ style to the new code. This
also applies to KeyValueBlobReader that was added recently.

Microbenchmark result below:

```
---------------------------------------------------------------------
Benchmark                           Time             CPU   Iterations
---------------------------------------------------------------------
BM_ParselessLMOpenClose         17710 ns        17199 ns        33422
BM_FastLMOpenClose          376520248 ns    367526500 ns            2
BM_ParselessLMFindUnigrams       5967 ns         5899 ns       113729
BM_FastLMFindUnigrams            2268 ns         2265 ns       307038
```

2022-01-15 16:15:02 -08:00

1 Commits