impala

mirror of https://github.com/apache/impala.git synced 2026-02-03 09:00:39 -05:00

Files

Michael Smith 1b6011c6a0 Revert "IMPALA-11253: Support testing with Java 11"

This reverts commit ee6395db76 as it is
not flexible enough at detecting Java automatically in likely build
environments.

Change-Id: I836c9f7fd10740b15f7e40b2e7f889ac7ee61fc3
Reviewed-on: http://gerrit.cloudera.org:8080/19908
Tested-by: Impala Public Jenkins <impala-public-jenkins@cloudera.com>
Reviewed-by: Michael Smith <michael.smith@cloudera.com>

2023-05-21 14:00:14 +00:00

src/main/java/org/apache/impala/infra/tableflattener

IMPALA-12077: Remove deprecated Avro methods

2023-04-21 22:53:46 +00:00

.gitignore

IMPALA-10198 (part 1): Unify Java in a single java/ directory

2020-10-15 19:30:13 +00:00

pom.xml

Revert "IMPALA-11253: Support testing with Java 11"

2023-05-21 14:00:14 +00:00

README

IMPALA-10198 (part 1): Unify Java in a single java/ directory

2020-10-15 19:30:13 +00:00

README

This is a tool to convert a nested dataset to an unnested dataset. The source and/or
destination can be the local file system or HDFS.

Structs get converted to a column (with a long name). Arrays and Maps get converted to
a table which can be joined with the parent table on id column.

$ mvn exec:java \
    -Dexec.mainClass=org.apache.impala.infra.tableflattener.Main \
    -Dexec.arguments="file:///tmp/in.parquet,file:///tmp/out,-sfile:///tmp/in.avsc"

$ mvn exec:java \
    -Dexec.mainClass=org.apache.impala.infra.tableflattener.Main \
    -Dexec.arguments="hdfs://localhost:20500/nested.avro,file://$PWD/unnested"

There are various options to specify the type of input file but the output is always
parquet/snappy.

For additional help, use the following command:
$ mvn exec:java \
    -Dexec.mainClass=org.apache.impala.infra.tableflattener.Main -Dexec.arguments="--help"

This is used by testdata/bin/generate-load-nested.sh.