Skip to content

GH-2870: Report compressed dictionary page sizes in the CLI - #3798

Merged
wgtmac merged 4 commits into
apache:masterfrom
1fanwang:1fannnw/fix-dictionary-page-size
Oct 11, 2026
Merged

wgtmac merged 4 commits into
apache:masterfrom
1fanwang:1fannnw/fix-dictionary-page-size

Conversation

@1fanwang

@1fanwang 1fanwang commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Rationale for this change

parquet pages reports decoded dictionary sizes instead of the bytes stored in compressed files. It can also fail on a valid column chunk whose dictionary page is present but unused by its data pages.

What changes are included in this PR?

The CLI reads compressed sizes from each stored page header and starts at the column chunk's physical first page.

Closes #2870.

Are these changes tested?

Testing Done

I built the baseline and fixed runtime jars with JDK 17.0.5 and Thrift 0.24.0.

$ runtime=parquet-cli/target/parquet-cli-1.19.0-SNAPSHOT-runtime.jar
$ git worktree add --detach ../parquet-page-size-base 2df8d02678dab4bb8b926a0d3221cc652984c7ab
$ (cd ../parquet-page-size-base && ./mvnw -B -ntp -pl parquet-cli -am -Plocal -DskipTests package)
$ cp "../parquet-page-size-base/$runtime" before-cli.jar
$ ./mvnw -B -ntp -pl parquet-cli -am -Plocal '-Dtest=ShowPagesCommandTest,ConvertCSVCommandTest' -Dsurefire.failIfNoSpecifiedTests=false package
$ cp "$runtime" after-cli.jar

$ python3 -c 'from pathlib import Path; Path("dictionary_page_input.csv").write_text("color\n" + ("a" * 120 + "\n" + "b" * 120 + "\n") * 100)'
$ java -Xmx512m -XX:ActiveProcessorCount=2 -jar before-cli.jar convert-csv dictionary_page_input.csv --require color --compression-codec GZIP -o compressed-dictionary.parquet

Compressed dictionary before:

$ java -Xmx512m -XX:ActiveProcessorCount=2 -jar before-cli.jar pages -c color compressed-dictionary.parquet
  0-D    dict  G _  2       124.00 B   248 B
  0-1    data  G R  200     0.13 B     25 B

Compressed dictionary after:

$ java -Xmx512m -XX:ActiveProcessorCount=2 -jar after-cli.jar pages -c color compressed-dictionary.parquet
  0-D    dict  G _  2       16.50 B    33 B
  0-1    data  G R  200     0.13 B     25 B

Save the reproducer source below as WriteUnusedDictionary.java, then create the unused-dictionary fixture.

Unused dictionary before:

$ java -Xmx512m -XX:ActiveProcessorCount=2 -cp before-cli.jar WriteUnusedDictionary.java unused-dictionary.parquet
$ java -Xmx512m -XX:ActiveProcessorCount=2 -jar before-cli.jar pages unused-dictionary.parquet
Unknown error
java.lang.RuntimeException: java.io.IOException: can not read class org.apache.parquet.format.PageHeader: Required field 'uncompressed_page_size' was not found in serialized data!

Unused dictionary after:

$ java -Xmx512m -XX:ActiveProcessorCount=2 -jar after-cli.jar pages unused-dictionary.parquet
  0-D    dict  _ _  3       4.00 B     12 B
  0-1    data  _ _  2       4.00 B     8 B

$ java -Xmx512m -XX:ActiveProcessorCount=2 -jar after-cli.jar cat unused-dictionary.parquet
{"value": 41}
{"value": 42}
Reproducer source: WriteUnusedDictionary.java
import java.util.Map;
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.parquet.bytes.BytesInput;
import org.apache.parquet.column.Encoding;
import org.apache.parquet.column.page.DictionaryPage;
import org.apache.parquet.column.statistics.Statistics;
import org.apache.parquet.hadoop.ParquetFileWriter;
import org.apache.parquet.hadoop.ParquetWriter;
import org.apache.parquet.hadoop.metadata.CompressionCodecName;
import org.apache.parquet.hadoop.util.HadoopOutputFile;
import org.apache.parquet.schema.MessageType;
import org.apache.parquet.schema.PrimitiveType;
import org.apache.parquet.schema.Types;

public class WriteUnusedDictionary {
  public static void main(String[] args) throws Exception {
    PrimitiveType type = Types.required(PrimitiveType.PrimitiveTypeName.INT32).named("value");
    MessageType schema = new MessageType("record", type);
    try (ParquetFileWriter writer = new ParquetFileWriter(
        HadoopOutputFile.fromPath(new Path(args[0]), new Configuration()), schema,
        ParquetFileWriter.Mode.CREATE, ParquetWriter.DEFAULT_BLOCK_SIZE,
        ParquetWriter.MAX_PADDING_SIZE_DEFAULT)) {
      writer.start();
      writer.startBlock(2);
      writer.startColumn(schema.getColumnDescription(new String[] {"value"}), 2, CompressionCodecName.UNCOMPRESSED);
      writer.writeDictionaryPage(new DictionaryPage(
          BytesInput.concat(BytesInput.fromInt(10), BytesInput.fromInt(20), BytesInput.fromInt(30)), 3, Encoding.PLAIN));
      writer.writeDataPage(2, 2 * Integer.BYTES,
          BytesInput.concat(BytesInput.fromInt(41), BytesInput.fromInt(42)),
          Statistics.createStats(type), Encoding.RLE, Encoding.RLE, Encoding.PLAIN);
      writer.endColumn();
      writer.endBlock();
      writer.end(Map.of());
    }
  }
}

Are there any user-facing changes?

Dictionary summaries report stored compressed bytes, including when no data page uses the dictionary.

Signed-off-by: 1fanwang <1fannnw@gmail.com>

@divjotarora divjotarora left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Seems reasonable, but one possible edge case

// TODO: the compressed size of a dictionary page is lost in Parquet
dict.getUncompressedSize();
long totalSize = dict.getCompressedSize();
long totalSize = getPageCompressedSize();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

getPageCompressedSize calls columnChunk.hasDictionaryPage() to determine the page offset. This function does:

  public boolean hasDictionaryPage() {
    EncodingStats stats = getEncodingStats();
    if (stats != null) {
      // ensure there is a dictionary page and that it is used to encode data pages
      return stats.hasDictionaryPages() && stats.hasDictionaryEncodedPages();
    }

    Set<Encoding> encodings = getEncodings();
    return (encodings.contains(PLAIN_DICTIONARY) || encodings.contains(RLE_DICTIONARY));
  }

In an edge case where a writer emits a dictionary page followed by no PLAIN_DICTIONARY or RLE_DICTIONARY data pages, we would get wrong results because hasDictionaryPage() returns false.

Realistically, writers wouldn't do this, but I didn't find any wording in the spec explicitly disallowing it. The TestParquetFileWriter code in this repo seems to do exactly this.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At least this change is not worse than before. :)

Signed-off-by: 1fanwang <1fannnw@gmail.com>
Dictionary presence and dictionary use are separate metadata facts. This regression keeps the CLI anchored to the physical chunk start for valid files with an unused dictionary.

Signed-off-by: 1fanwang <1fannnw@gmail.com>
…ary-page-size

Signed-off-by: 1fanwang <1fannnw@gmail.com>
@wgtmac
wgtmac merged commit 47382e6 into apache:master Oct 11, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The page compressedSize printed by the ShowPagesCommand is actually uncompressedSize

3 participants