PostgreSQL: penggunaan kembali hasil antara yang kompleks dalam permintaan yang sama

8

Menggunakan PostgreSQL (8,4), saya menciptakan pandangan yang merangkum berbagai hasil dari beberapa tabel (misalnya membuat kolom a, b, cdalam tampilan), dan kemudian saya harus menggabungkan beberapa hasil ini bersama-sama dalam query yang sama (misalnya a+b, a-b, (a+b)/c, ...), sehingga menghasilkan hasil akhir. Yang saya perhatikan adalah bahwa hasil antara sepenuhnya dihitung setiap kali digunakan, bahkan jika itu dilakukan dalam permintaan yang sama.

Apakah ada cara untuk mengoptimalkan ini untuk menghindari hasil yang sama yang akan dihitung setiap waktu?

Berikut adalah contoh sederhana yang mereproduksi masalah.

CREATE TABLE test1 (
    id SERIAL PRIMARY KEY,
    log_timestamp TIMESTAMP NOT NULL
);
CREATE TABLE test2 (
    test1_id INTEGER NOT NULL REFERENCES test1(id),
    category VARCHAR(10) NOT NULL,
    col1 INTEGER,
    col2 INTEGER
);
CREATE INDEX test_category_idx ON test2(category);

-- Added after edit to this question
CREATE INDEX test_id_idx ON test2(test1_id);

-- Populating with test data.
INSERT INTO test1(log_timestamp)
    SELECT * FROM generate_series('2011-01-01'::timestamp, '2012-01-01'::timestamp, '1 hour');
INSERT INTO test2
    SELECT id, substr(upper(md5(random()::TEXT)), 1, 1),
               (20000*random()-10000)::int, (3000*random()-200)::int FROM test1;
INSERT INTO test2
    SELECT id, substr(upper(md5(random()::TEXT)), 1, 1),
               (2000*random()-1000)::int, (3000*random()-200)::int FROM test1;
INSERT INTO test2
    SELECT id, substr(upper(md5(random()::TEXT)), 1, 1),
               (2000*random()-40)::int, (3000*random()-200)::int FROM test1;

Berikut ini adalah tampilan yang melakukan operasi yang paling memakan waktu:

CREATE VIEW testview1 AS
    SELECT
       t1.id,
       t1.log_timestamp,
       (SELECT SUM(t2.col1) FROM test2 t2 WHERE t2.test1_id=t1.id AND category='A') AS a,
       (SELECT SUM(t2.col2) FROM test2 t2 WHERE t2.test1_id=t1.id AND category='B') AS b,
       (SELECT SUM(t2.col1 - t2.col2) FROM test2 t2 WHERE t2.test1_id=t1.id AND category='C') AS c
    FROM test1 t1;

SELECT a FROM testview1menghasilkan rencana ini (melalui EXPLAIN ANALYZE):

 Seq Scan on test1 t1  (cost=0.00..1787086.55 rows=8761 width=4) (actual time=12.877..10517.575 rows=8761 loops=1)
   SubPlan 1
     ->  Aggregate  (cost=203.96..203.97 rows=1 width=4) (actual time=1.193..1.193 rows=1 loops=8761)
           ->  Bitmap Heap Scan on test2 t2  (cost=36.49..203.95 rows=1 width=4) (actual time=1.109..1.177 rows=0 loops=8761)
                 Recheck Cond: ((category)::text = 'A'::text)
                 Filter: (test1_id = $0)
                 ->  Bitmap Index Scan on test_category_idx  (cost=0.00..36.49 rows=1631 width=0) (actual time=0.414..0.414 rows=1631 loops=8761)
                       Index Cond: ((category)::text = 'A'::text)
 Total runtime: 10522.346 ms

SELECT a, a FROM testview1menghasilkan rencana ini :

 Seq Scan on test1 t1  (cost=0.00..3574037.50 rows=8761 width=4) (actual time=3.343..20550.817 rows=8761 loops=1)
   SubPlan 1
     ->  Aggregate  (cost=203.96..203.97 rows=1 width=4) (actual time=1.183..1.183 rows=1 loops=8761)
           ->  Bitmap Heap Scan on test2 t2  (cost=36.49..203.95 rows=1 width=4) (actual time=1.100..1.166 rows=0 loops=8761)
                 Recheck Cond: ((category)::text = 'A'::text)
                 Filter: (test1_id = $0)
                 ->  Bitmap Index Scan on test_category_idx  (cost=0.00..36.49 rows=1631 width=0) (actual time=0.418..0.418 rows=1631 loops=8761)
                       Index Cond: ((category)::text = 'A'::text)
   SubPlan 2
     ->  Aggregate  (cost=203.96..203.97 rows=1 width=4) (actual time=1.154..1.154 rows=1 loops=8761)
           ->  Bitmap Heap Scan on test2 t2  (cost=36.49..203.95 rows=1 width=4) (actual time=1.083..1.143 rows=0 loops=8761)
                 Recheck Cond: ((category)::text = 'A'::text)
                 Filter: (test1_id = $0)
                 ->  Bitmap Index Scan on test_category_idx  (cost=0.00..36.49 rows=1631 width=0) (actual time=0.426..0.426 rows=1631 loops=8761)
                       Index Cond: ((category)::text = 'A'::text)
 Total runtime: 20557.581 ms

Di sini, memilih a, amemakan waktu dua kali lebih lama daripada memilih a, sedangkan mereka benar-benar dapat dihitung hanya sekali. Sebagai contoh, dengan SELECT a, a+b, a-b FROM testview1, melewati sub-rencana untuk a3 kali dan melalui bdua kali, sedangkan waktu eksekusi dapat dikurangi menjadi 2/5 dari total waktu (dengan asumsi + dan - dapat diabaikan di sini).

Ini adalah hal yang baik karena tidak menghitung kolom yang tidak digunakan ( bdan c) ketika tidak diperlukan, tetapi apakah ada cara untuk membuatnya menghitung kolom yang digunakan sama dari tampilan hanya sekali?

EDIT: @ Frank Heikens dengan benar menyarankan untuk menggunakan indeks, yang tidak ada dalam contoh di atas. Meskipun meningkatkan kecepatan untuk setiap sub-paket, itu tidak mencegah sub-query yang sama untuk dihitung berulang kali. Maaf, saya harus meletakkan ini di pertanyaan awal untuk membuatnya lebih jelas.

performance postgresql Bruno
sumber

6

(Permintaan maaf untuk menjawab pertanyaan saya sendiri, tetapi setelah membaca pertanyaan dan jawaban yang tidak terkait ini , terlintas di benak saya, saya harus mencoba menggunakan CTE sebagai gantinya. Ini berhasil.)

Berikut ini adalah pandangan lain, mirip dengan testview1dalam pertanyaan, tetapi yang menggunakan Ekspresi Tabel Umum :

CREATE VIEW testview2 AS
    WITH testcte AS (SELECT
       t1.id,
       t1.log_timestamp,
       (SELECT SUM(t2.col1) FROM test2 t2 WHERE t2.test1_id=t1.id AND category='A') AS a,
       (SELECT SUM(t2.col2) FROM test2 t2 WHERE t2.test1_id=t1.id AND category='B') AS b,
       (SELECT SUM(t2.col1 - t2.col2) FROM test2 t2 WHERE t2.test1_id=t1.id AND category='C') AS c
      FROM test1 t1)
    SELECT * FROM testcte;

(Ini hanya sebuah contoh, saya tidak menyarankan bahwa menggabungkan pandangan dan CTE adalah ide yang bagus: CTE mungkin cukup.)

Tidak seperti testview1, rencana kueri untuk SELECT a FROM testview2saat ini juga menghitung bdan c, yang diabaikan sejak tidak digunakan di testview1:

Subquery Scan testview2  (cost=395272.42..395535.25 rows=8761 width=8) (actual time=0.256..607.941 rows=8761 loops=1)
   ->  CTE Scan on testcte  (cost=395272.42..395447.64 rows=8761 width=36) (actual time=0.255..604.106 rows=8761 loops=1)
         CTE testcte
           ->  Seq Scan on test1 t1  (cost=0.00..395272.42 rows=8761 width=12) (actual time=0.252..589.358 rows=8761 loops=1)
                 SubPlan 1
                   ->  Aggregate  (cost=15.02..15.03 rows=1 width=4) (actual time=0.021..0.021 rows=1 loops=8761)
                         ->  Bitmap Heap Scan on test2 t2  (cost=4.28..15.02 rows=1 width=4) (actual time=0.015..0.015 rows=0 loops=8761)
                               Recheck Cond: (test1_id = $0)
                               Filter: ((category)::text = 'A'::text)
                               ->  Bitmap Index Scan on test_if_idx  (cost=0.00..4.28 rows=3 width=0) (actual time=0.009..0.009 rows=3 loops=8761)
                                     Index Cond: (test1_id = $0)
                 SubPlan 2
                   ->  Aggregate  (cost=15.02..15.03 rows=1 width=4) (actual time=0.019..0.019 rows=1 loops=8761)
                         ->  Bitmap Heap Scan on test2 t2  (cost=4.28..15.02 rows=1 width=4) (actual time=0.012..0.012 rows=0 loops=8761)
                               Recheck Cond: (test1_id = $0)
                               Filter: ((category)::text = 'B'::text)
                               ->  Bitmap Index Scan on test_if_idx  (cost=0.00..4.28 rows=3 width=0) (actual time=0.007..0.007 rows=3 loops=8761)
                                     Index Cond: (test1_id = $0)
                 SubPlan 3
                   ->  Aggregate  (cost=15.02..15.04 rows=1 width=8) (actual time=0.020..0.020 rows=1 loops=8761)
                         ->  Bitmap Heap Scan on test2 t2  (cost=4.28..15.02 rows=1 width=8) (actual time=0.013..0.014 rows=0 loops=8761)
                               Recheck Cond: (test1_id = $0)
                               Filter: ((category)::text = 'C'::text)
                               ->  Bitmap Index Scan on test_if_idx  (cost=0.00..4.28 rows=3 width=0) (actual time=0.007..0.007 rows=3 loops=8761)
                                     Index Cond: (test1_id = $0)

Namun, itu tidak menghitung ulang hasil yang digunakan beberapa kali dalam kueri yang sama (yang merupakan tujuan).

Berbeda testview1dengan yang SELECT a, a, a, a, amembutuhkan waktu 5 kali lebih lama daripada SELECT a, di sini SELECT a, a, a, a, a, b, c, a+b, a+c, b+c FROM testview2hanya selama SELECT a FROM testview2atau SELECT a, b, c FROM testview2. Hanya melewati a, bdan csekali:

 Subquery Scan testview2  (cost=395272.42..395600.96 rows=8761 width=24) (actual time=0.147..562.790 rows=8761 loops=1)
   ->  CTE Scan on testcte  (cost=395272.42..395447.64 rows=8761 width=36) (actual time=0.144..554.194 rows=8761 loops=1)
         CTE testcte
           ->  Seq Scan on test1 t1  (cost=0.00..395272.42 rows=8761 width=12) (actual time=0.140..542.657 rows=8761 loops=1)
                 SubPlan 1
                   ->  Aggregate  (cost=15.02..15.03 rows=1 width=4) (actual time=0.019..0.019 rows=1 loops=8761)
                         ->  Bitmap Heap Scan on test2 t2  (cost=4.28..15.02 rows=1 width=4) (actual time=0.012..0.013 rows=0 loops=8761)
                               Recheck Cond: (test1_id = $0)
                               Filter: ((category)::text = 'A'::text)
                               ->  Bitmap Index Scan on test_if_idx  (cost=0.00..4.28 rows=3 width=0) (actual time=0.007..0.007 rows=3 loops=8761)
                                     Index Cond: (test1_id = $0)
                 SubPlan 2
                   ->  Aggregate  (cost=15.02..15.03 rows=1 width=4) (actual time=0.019..0.019 rows=1 loops=8761)
                         ->  Bitmap Heap Scan on test2 t2  (cost=4.28..15.02 rows=1 width=4) (actual time=0.012..0.012 rows=0 loops=8761)
                               Recheck Cond: (test1_id = $0)
                               Filter: ((category)::text = 'B'::text)
                               ->  Bitmap Index Scan on test_if_idx  (cost=0.00..4.28 rows=3 width=0) (actual time=0.006..0.006 rows=3 loops=8761)
                                     Index Cond: (test1_id = $0)
                 SubPlan 3
                   ->  Aggregate  (cost=15.02..15.04 rows=1 width=8) (actual time=0.018..0.019 rows=1 loops=8761)
                         ->  Bitmap Heap Scan on test2 t2  (cost=4.28..15.02 rows=1 width=8) (actual time=0.012..0.012 rows=0 loops=8761)
                               Recheck Cond: (test1_id = $0)
                               Filter: ((category)::text = 'C'::text)
                               ->  Bitmap Index Scan on test_if_idx  (cost=0.00..4.28 rows=3 width=0) (actual time=0.007..0.007 rows=3 loops=8761)
                                     Index Cond: (test1_id = $0)

Bruno
sumber

1

Tidak perlu minta maaf! :-) Sebenarnya ada lencana Self Learner untuk menjawab pertanyaan Anda sendiri.

Gayus

3

Anda memerlukan indeks pada test1_id dalam tabel test2, itu akan mengubah banyak hal.

Seq Scan on test1 t1  (cost=0.00..301450.63 rows=8761 width=12) (actual time=0.108..229.859 rows=8761 loops=1)
  SubPlan 1
    ->  Aggregate  (cost=11.45..11.46 rows=1 width=4) (actual time=0.008..0.008 rows=1 loops=8761)
          ->  Bitmap Heap Scan on test2 t2  (cost=3.27..11.45 rows=1 width=4) (actual time=0.007..0.007 rows=0 loops=8761)
                Recheck Cond: (test1_id = t1.id)
                Filter: ((category)::text = 'A'::text)
                ->  Bitmap Index Scan on idx_id  (cost=0.00..3.27 rows=3 width=0) (actual time=0.003..0.003 rows=3 loops=8761)
                      Index Cond: (test1_id = t1.id)
  SubPlan 2
    ->  Aggregate  (cost=11.45..11.46 rows=1 width=4) (actual time=0.007..0.008 rows=1 loops=8761)
          ->  Bitmap Heap Scan on test2 t2  (cost=3.27..11.45 rows=1 width=4) (actual time=0.006..0.006 rows=0 loops=8761)
                Recheck Cond: (test1_id = t1.id)
                Filter: ((category)::text = 'B'::text)
                ->  Bitmap Index Scan on idx_id  (cost=0.00..3.27 rows=3 width=0) (actual time=0.003..0.003 rows=3 loops=8761)
                      Index Cond: (test1_id = t1.id)
  SubPlan 3
    ->  Aggregate  (cost=11.46..11.47 rows=1 width=8) (actual time=0.008..0.008 rows=1 loops=8761)
          ->  Bitmap Heap Scan on test2 t2  (cost=3.27..11.45 rows=1 width=8) (actual time=0.006..0.006 rows=0 loops=8761)
                Recheck Cond: (test1_id = t1.id)
                Filter: ((category)::text = 'C'::text)
                ->  Bitmap Index Scan on idx_id  (cost=0.00..3.27 rows=3 width=0) (actual time=0.003..0.003 rows=3 loops=8761)
                      Index Cond: (test1_id = t1.id)
Total runtime: 232.419 ms

Frank Heikens
sumber

Terima kasih, ini memang meningkatkan kinerja, tetapi tabel saya sebenarnya sudah memiliki indeks yang sesuai sebenarnya. Masalah awal masih ada: SELECT a, a, a, a, a FROM testview1masih membutuhkan waktu 5 kali lebih lama dari SELECT a FROM testview1.

Bruno

PostgreSQL: penggunaan kembali hasil antara yang kompleks dalam permintaan yang sama

Jawaban: