Posts

Showing posts with the label Oracle

Using DBGen to Generate TPC-H Test Data

Image
In this post I describe the Data Base Generator (DBGen) utility in terms of what it is, how to install it, and how to use it.   This post assumes a Linux operating system.   There are different build procedures for DBGen on Windows, but the concepts carry over. DBGen is used to create a TPC-H schema (i.e., it provides the CREATE TABLE statements) and to generate the data.   DBGen is used in official TPC-H benchmarking, but it can also be used in informal TPC-H-like benchmarking.   The schema and data is compatible with most 3 rd party benchmarking tools like HammerDB, so even if you plan to run a TPC-H-like test using HammerDB you may still use DBGen to create the schema and data. The  CREATE TABLE   statements are ANSI standard SQL and can be run in any DBMS without editing.   There are no SQL statements for indexes or foreign key constraints provided by DBGen since those types of objects are not required.   The data is created as delimited...

Understanding Database Latency

Image
Latency is basically the amount of time you must wait to get a response from a request.   It is measured at a specific component like a storage or network device.   When you ask a computer to do something, each component involved in that request has a minimum amount of time it takes to reply with an answer even if the answer is a null value.   If your database requests a single block from storage, then the time it takes to receive that block is storage latency.   If the storage is attached over a network like Fibre Channel or SAS, then the network interfaces and cables each add their own latency. Latencies are often measured in milliseconds (ms).   There are 1,000 milliseconds per second, and most people cannot perceive anything smaller than 1 millisecond.   However, computers are wicked fast and the latency may need to be measured in microseconds (each μs is a millionth of a second) or even nanoseconds (each ns is a billionth of a second).   Hard dr...

Ignoring Rules When Benchmarking Databases

This post is part of a series about benchmarking the performance of relational database.  My earlier posts detailed the official TPC benchmarks for relational databases and the popular tools for running informal versions of those benchmarks.  This post describes some of the rules that drive cost and complexity, which leads to the most commonly ignored rules. Each database benchmark has a strict set of rules called a "specification".   There is one specification for TPC-C, another for TPC-E, and so on for every TPC benchmark.   Following all of the rules in a given specification is expensive and time consuming.   Some of the rules aren't practical and can be ignored when testing informally or internally.   You are only required to follow the rules when testing with the intent to publish or if you plan to use the TPC trademark in a publication. To be honest, I have never performed a formal TPC benchmark or followed all of the rules.   All of my tests...

TPC-like Database Benchmarking Tools

Image
  This post describes five popular tools for benchmarking relational databases: HammerDB, Quest Benchmark Factory, SwingBench, pgbench, and Oracle RAT.  Previously, I described TPC and TPC-like benchmarks for relational databases.   Recall a benchmark is "like" a TPC benchmark if it follows some of the rules but not all of them.   The tools described in this post follow some of the TPC's rules so we call them "TPC-like".   Publishing results is generally not allowed by commercial DBMS vendors, and use of the benchmark trademark is also restricted. Ok then, let's look at the most commonly used benchmarking tools for relational databases. HammerDB is my #1 choice for benchmarking relational databases.   HammerDB supports every major DBMS on Linux and Windows including bare metal, VM, and cloud.   This means you can use one tool for all environments.   You can even compare performance across platforms or DBMS, such as comparing Oracle on Linux...