impala

mirror of https://github.com/apache/impala.git synced 2026-01-03 06:00:52 -05:00

Author	SHA1	Message	Date
Lars Volker	b5570da405	IMPALA-1740: Add support for skip.header.line.count. HIVE-5795 introduced a parameter skip.header.line.count to skip header lines from input files. This change introduces the capability to skip an arbitrary number of header lines from csv input files on hdfs. The size of the total file header must be smaller than max_scan_range_length, otherwise an error will be reported. This is necessary because scan ranges are not read in disk order, so there is no way of identifying header lines except by counting from the start of the first scan range. [localhost:21000] > alter table t1 set tblproperties('skip.header.line.count'='1'); Query: alter table t1 set tblproperties('skip.header.line.count'='1') [localhost:21000] > select * from t1; Query: select * from t1 +----+----+ \| c1 \| c2 \| +----+----+ \| 1 \| 1 \| \| 2 \| 2 \| \| 3 \| 3 \| +----+----+ Fetched 3 row(s) in 0.32s [localhost:21000] > alter table t1 set tblproperties('skip.header.line.count'='0'); Query: alter table t1 set tblproperties('skip.header.line.count'='0') [localhost:21000] > select * from t1; Query: select * from t1 +------+------+ \| c1 \| c2 \| +------+------+ \| NULL \| NULL \| \| 1 \| 1 \| \| 2 \| 2 \| \| 3 \| 3 \| +------+------+ WARNINGS: Error converting column: 0 TO INT (Data is: num1) Error converting column: 1 TO DOUBLE (Data is: num2) file: hdfs://localhost:20500/test-warehouse/t1/test.txt record: num1,num2 Fetched 4 row(s) in 0.41s Change-Id: I595f01a165d41499ca1956fe748ba3840a6eb543 Reviewed-on: http://gerrit.cloudera.org:8080/2110 Reviewed-by: Lars Volker <lv@cloudera.com> Tested-by: Internal Jenkins	2016-05-12 14:17:46 -07:00
Jim Apple	83a1434bc9	IMPALA-2840: Don't store table location in partition location For a table with location "ABC", most partitions will have locations like "ABC/DEF=2". The "ABC" part of the location does not need to be stored in Catalog for each partition; we can compress it down to one int in the common case. This is done by stripping from each partition location the last N directories (where N is the number of clustering columns) and storing the resulting string in a cache of partition location prefixes. In the cache, this location prefix string is mapped to an int. Partition locations are then stored as a tuple consisting of that int and a suffix string; the partition location can be reconstructed as the concatenation of the prefix string (from the cache) and the suffix. Though this scheme was designed in the expectation that most partitions will be stored in directories like "/part_col_1=1.23/part_col_2=234/", it works even when that is not the case. TODO: Since each partition stores the literal values for the partitioning columns, we could also elide the column names and values when partitions are placed in directories like "/part_col_1=1.23/part_col_2=234/" Change-Id: I8c67b6ce0f83de2f5277a528a9ce67e47d638adb Reviewed-on: http://gerrit.cloudera.org:8080/2355 Reviewed-by: Jim Apple <jbapple@cloudera.com> Tested-by: Internal Jenkins	2016-05-12 14:17:29 -07:00
Alex Behm	5c0e1fa1e8	IMPALA-2974: Use Type.toSql() instead of toString() in ALTER TABLE CHANGE COLUMN. Change-Id: I140bdea755e44d3f2ceb4a8f5e288faaddaa963f Reviewed-on: http://gerrit.cloudera.org:8080/2285 Reviewed-by: Alex Behm <alex.behm@cloudera.com> Tested-by: Internal Jenkins	2016-02-26 15:37:24 -08:00
Dimitris Tsirogiannis	b78b37cbdc	IMPALA-1480: Slow DDL statements with large number of partitions This commit improves the performance of DDL statements on tables with large number of partitions. Previously, the catalog would force-reload the entire table metadata during the execution of DDL and insert statements, causing significant delays for tables with large number of partitions. With this commit the catalog is reusing any cached table entries to partially reload table metadata for only those partitions that have been modified. With this change we've improved the performance of some DDL and insert statements by at least 4-5X. This commit also adds basic table-level locking to protect table metadata from concurrent DDL operations. Preliminary performance measurements ----------------------------------- Workload: insert into table partition () select ... limit 10 Iterations: 10 Num partitions OLD (avg time sec) NEW (avg time sec) 1K 1.15 0.45 5K 3.65 0.9 10K 5.75 1.38 15K 10.1 2.02 30K 25.4 4.46 Workload: alter table partition() set location... Iterations: 10 Num partitions OLD (avg time sec) NEW (avg time sec) 1K 0.8 0.47 5K 4.3 0.71 10K 7.1 1.2 15K 13.2 1.8 30K 26.8 3.4 Change-Id: I4da7fb6df0a71162b0cb60e6025a4019cb9572bf Reviewed-on: http://gerrit.cloudera.org:8080/1706 Reviewed-by: Dimitris Tsirogiannis <dtsirogiannis@cloudera.com> Tested-by: Internal Jenkins	2016-01-28 09:18:57 +00:00
Tim Armstrong	a280b93a37	IMPALA-2070: include the database comment when showing databases As part of change, refactor catalog and frontend functions to return TDatabase/Db objects instead of just the string names of databases - this required a lot of method/variable renamings. Add test for creating database with comment. Modify existing tests that assumed only a single column in SHOW DATABASES results. Change-Id: I400e99b0aa60df24e7f051040074e2ab184163bf Reviewed-on: http://gerrit.cloudera.org:8080/620 Reviewed-by: Tim Armstrong <tarmstrong@cloudera.com> Tested-by: Internal Jenkins	2015-12-04 01:40:41 +00:00
Jim Apple	19b6bf0201	IMPALA-2226: Throw AnalysisError if table properties are too large (for the Hive metastore) This only enforced the defaults in Hive. Users how manually choose to change the schema in Hive may trigger these new analysis exceptions in this commit unnecessarily. The Hive issue tracking the length restrictions is https://issues.apache.org/jira/browse/HIVE-9815 Change-Id: Ia30f286193fe63e51a10f0c19f12b848c4b02f34 Reviewed-on: http://gerrit.cloudera.org:8080/721 Reviewed-by: Jim Apple <jbapple@cloudera.com> Tested-by: Internal Jenkins	2015-10-29 19:39:25 +00:00
Alex Behm	14a8cadcf6	Nested Types: Pretty print complex types in DESCRIBE. The current DESCRIBE prints the column type as a single string without whitespace. As a result, the DESCRIBE output for tables with complex types is basically unreadable/unusable, e.g., from the Impala shell. This patch adds a prettyPrint() function to the FE Type and uses that for generating a nicely formatted DESCRIBE output. The output of DESCRIBE FORMATTED is intentionally not modified because exact Hive-compatibility has been and presumably continues to be very important to our users. Change-Id: Ida810facdffd970948b837b83a60f9ddcd95f44d Reviewed-on: http://gerrit.cloudera.org:8080/633 Reviewed-by: Alex Behm <alex.behm@cloudera.com> Tested-by: Internal Jenkins	2015-08-22 09:26:35 +00:00
Henry Robinson	f22b8659fd	IMPALA-1595: Add 'location' to SHOW [TABLE STATS\|PARTITIONS] for HDFS tables This patch adds a 'location' column to the output of SHOW TABLE STATS / SHOW PARTITIONS. This helps users understand the effects of ALTER TABLE SET LOCATION commands, particularly for partitions, and is easier to identify than the output of DESCRIBE FORMATTED. Some existing tests in alter-table.test have been updated to include checking the location output before and after a SET LOCATION command. The tests in show.test have also been updated to check for the location; all other tests that use SHOW [TABLE STATS\|PARTITIONS] use a generic regex to avoid overly verbose tests. Change-Id: I9d276f7b133c38c9319e0906397ca1c31cec95bb Reviewed-on: http://gerrit.cloudera.org:8080/316 Reviewed-by: Henry Robinson <henry@cloudera.com> Tested-by: Internal Jenkins	2015-04-21 19:27:50 +00:00
Dan Hecht	c8fb10f50a	S3: Some more work toward enabling additional S3 test coverage Add skip markers for S3 that can be used to categorize the tests that are skipped against S3 to help see what coverage is missing. Soon we'll be reworking some tests and/or adding new tests to get back the important gaps. Also, add a mechanism to parameterize paths in the .test files, and start using these new variables. This is a step toward enabling some more tests against S3. Finally, a fix for buildall.sh to stop the minicluster before applying the metastore snapshot. Otherwise, this fails since the ms db is in use. Change-Id: I142434ed67bed407e61d7b2c90f825734fc0dce0 Reviewed-on: http://gerrit.cloudera.org:8080/127 Reviewed-by: Dan Hecht <dhecht@cloudera.com> Tested-by: Internal Jenkins	2015-03-03 08:29:13 +00:00
Alex Behm	ffd124b48e	IMPALA-1711: Save, drop and restore col stats when renaming a table across dbs. The HMS does not properly migrate column stats when moving a table across databases (HIVE-9720). Apart from losing the stats, this issue also prevents the newly renamed table from being dropped. To work around the issue, this patch manually drops+adds the column stats when renaming a table across databases. Change-Id: If901c5d1e9a6b2cedc35034a537f18c361c8ffa1 Reviewed-on: http://gerrit.cloudera.org:8080/72 Reviewed-by: Alex Behm <alex.behm@cloudera.com> Tested-by: Internal Jenkins	2015-02-24 00:10:35 +00:00
Martin Grund	cee1e84c1e	IMPALA-1587: Extending caching directives for multiple replicas This patch adds the possibility to specify the number of replicas that should be cached in main memory. This can be useful in high QPS scenarios as the majority of the load is no longer the single cached replica, but a set of cached replicas. While the cache replication factor can be larger than the block replication factor on disk, the difference will be ignored by HDFS until more replicas become available. This extends the current syntax for specifying the cache pool in the following way: cached in 'poolName' is extended with the optional replication factor cached in 'poolName' with replication = XX By default, the cache replication factor is set to 1. As this value is not yet configurable in HDFS it's defined as a constant in the JniCatalog thrift specification. If a partitioned table is cached, all its child partitions inherit this cache replication factor. If child partitions have a custom cache replication factor, changing the cache replication factor on the partitioned table afterwards will overwrite this custom value. If a new partition is added to the table, it will again inherit the cache replication factor of the parent independent of the cache pool that is used to cache the partition. To review changes and status of the replication factor for tables and partitions the replication factor is part of output of the "show partitions" command. Change-Id: I2aee63258d6da14fb5ce68574c6b070cf948fb4d Reviewed-on: http://gerrit.sjc.cloudera.com:8080/5533 Tested-by: jenkins Reviewed-by: Martin Grund <mgrund@cloudera.com>	2015-01-26 20:30:59 -08:00
Henry Robinson	44f57e5fb6	IMPALA-1122: Compute stats with partition granularity This patch adds the ability to compute and drop column and table statistics at partition granularity. The following commands are added. Detail about the implementation follows. COMPUTE INCREMENTAL STATS <tbl_name> [PARTITION <partition_spec>] This variant of COMPUTE STATS will, ultimately, do the same thing as the traditional COMPUTE STATS statement, but does so by caching the intermediate state of the computation for each partition in the Hive MetaStore. If the PARTITION clause is added, the computation is performed for only that partition. If the PARTITION clause is omitted, incremental stats are updated only for those partitions with missing incremental stats (e.g. one column does not have stats, or incremental stats was never computed for this partition). In this patch, incremental stats are only invalidated when a DROP STATS variant is executed. Future patches can automatically invalidate the statistics after REFRESH or INSERT queries, etc. DROP INCREMENTAL STATS <tbl_name> PARTITION <part_spec> This variant of DROP stats removes the incremental statistics for the given table. It does not recalculate the statistics for the whole table, so this should be used only to invalidate the intermediate state for a partition which will shortly be subject to COMPUTE INCREMENTAL STATS. The point of this variant is to allow users to notify Impala when they believe a partition has changed significantly enough to warrant recomputation of its statistics. It is not necessary for new partitions; Impala will detect that they do not have any valid statistics. -------- This is achieved by adapting the existing HLL UDA via swapping its finalize method for a new one which returns the intermediate HLL buckets, rather than aggregating and then disposing of them. This intermediate state is then returned to Impala's catalog-op-executor.cc, which then passes the intermediate state back to the frontend to be ultimately stored in the HMS. This intermediate state is computed on a per-partition basis by grouping the input to the UDA by partition. Thus, the incremental computation produces one row for each partition selected (the set of which might be quite small, if there are few partitions without valid incremental stats: this is the point of the new commands). At the same time, the query coordinator aggregates the output of the UDA to produce table-level statistics. This computation incorporates any existing (and not re-computed) intermediate partition state which is passed to the coordinator by the frontend. The resulting statistics are saved to the table as normal. Intermediate statistics are serialised to the HMS by writing a Thrift structure's serialised form to the partition's 'parameters' map. There is a schema-imposed limit of 4000 characters to the serialised string, which is exacerbated by the fact that the Thrift representation must first be base-64 encoded to avoid type errors in the HMS. The current patch breaks the encoded structure into 4k chunks, and then recombines them on read. The alltypes table (11 columns) takes about three of these chunks. This may mean that incremental stats are not suitable for particularly wide tables: these structures could be zipped before encoding for some space savings. In the meantime, the NDV estimates are run-length encoded (since they are generally sparse); this can result in substantial space savings. Change-Id: If82cf4753d19eb532265acb556f798b95fbb0f34 Reviewed-on: http://gerrit.sjc.cloudera.com:8080/4475 Tested-by: jenkins Reviewed-by: Henry Robinson <henry@cloudera.com> Reviewed-on: http://gerrit.sjc.cloudera.com:8080/5408	2014-11-25 09:13:37 -08:00
Henry Robinson	6af7c8fe4a	IMPALA-1330: Fix column types for SHOW {table, partition} STATS Because we add 'total' to the last row in SHOW PARTITIONS, we set the partition key columns to be string. At least, that's what the comment said, but we didn't do that in fact. This patch also corrects the column type for max width, which should be INT. Change-Id: I787ab17be27f45107340119017e528c58a3daad3 Reviewed-on: http://gerrit.sjc.cloudera.com:8080/4678 Reviewed-by: Henry Robinson <henry@cloudera.com> Tested-by: jenkins	2014-10-06 15:16:56 -07:00
Lenni Kuff	286e312460	[CDH5] Minor code changes for Hive .13 support Changes include: * Fix compile errors due to new column stats API and other stats related fixes. * Temporarily disable JDBC tests due to new serialization format in Hive .13 * Disable view compatibility tests until we can get them to work in Hive .13 * Test fixes due to Hive's type checking for partition column values Change-Id: I05cc6a95976e0e037be79d91bc330a06d2fdc46c	2014-08-11 09:53:02 -07:00
Alex Behm	19bab59854	Create/alter/describe tables with complex types. This patch adds parsing of complex types and tests for using complex types in various exprs and create/alter/describe stmts. Change-Id: Ibc211a560c889f5ccfb616813700b923c89d8245 Reviewed-on: http://gerrit.ent.cloudera.com:8080/3577 Reviewed-by: Alex Behm <alex.behm@cloudera.com> Tested-by: jenkins Reviewed-on: http://gerrit.ent.cloudera.com:8080/3594	2014-07-23 17:26:14 -07:00
Ippokratis Pandis	e34ede292c	IMPALA-1016: Return correct number of NULL values when projecting newly added column This patch handles the case where when a query was projecting a newly added column, the parquet scanner was returning infinite values. Change-Id: Ie5f4d4a88d5868e8d9e5c39fa9440821776dde3c Reviewed-on: http://gerrit.ent.cloudera.com:8080/2725 Reviewed-by: Marcel Kornacker <marcel@cloudera.com> Tested-by: jenkins Reviewed-on: http://gerrit.ent.cloudera.com:8080/2761 Reviewed-by: Ippokratis Pandis <ipandis@cloudera.com>	2014-06-01 01:28:25 -07:00
Lenni Kuff	745c091fcc	[CDH5] Update SHOW TABLE STATS to include per-partition HDFS caching stats Change-Id: I71b01f84bbd308108d775e78c644e867b48e05be Reviewed-on: http://gerrit.ent.cloudera.com:8080/2621 Reviewed-by: Lenni Kuff <lskuff@cloudera.com> Tested-by: jenkins	2014-05-28 08:54:54 -07:00
Lenni Kuff	c45e9a70d9	[CDH5] Add DDL support for HDFS caching This change adds DDL support for HDFS caching. The DDL allows the user to indicate a table or partition should be cached and which pool to cache the data into: * Create a cached table: CREATE TABLE ... CACHED IN 'poolName' * Cache a table/partition: ALTER TABLE ... [partitionSpec] SET CACHED IN 'poolName' * Uncache a table/partition: ALTER TABLE ... [partitionSpec] SET UNCACHED When a table/partition is marked as cached, a new HDFS caching request is submitted to cache the location (HDFS path) of the table/partition and the ID of that request is stored with in the table metadata (in the table properties). This is stored as: 'cache_directive_id'='<requestId>'. The cache requests and IDs are managed by HDFS and persisted across HDFS restarts. When a cached table or partition is dropped it is important to uncache the cached data (drop the associated cache request). For partitioned tables, this means dropping all cache requests from all cached partitions in the table. Likewise, if a partitioned table is created as cached, new partitions should be marked as cached by default. It is desirable to know which cache pools exists early on (in analysis) so the query will fail without hitting HDFS/CatalogServer if a non-existent pool is specified. To support this, a new cache pool catalog object type was introduced. The catalog server caches the known pools (periodically refreshing the cache) and sends the known pools out in catalog updates. This allows impalads to perform analysis checks on cache pool existence going to HDFS. It would be easy to use this to add basic cache pool management in the future (ADD/DROP/SHOW CACHE POOL). Waiting for the table/partition to become cached may take a long time. Instead of blocking the user from access the time during this period we will wait for the cache requests to complete in the background and once they have finished the table metadata will be automatically refreshed. Change-Id: I1de9c6e25b2a3bdc09edebda5510206eda3dd89b Reviewed-on: http://gerrit.ent.cloudera.com:8080/2310 Reviewed-by: Lenni Kuff <lskuff@cloudera.com> Tested-by: jenkins	2014-05-27 16:47:15 -07:00
Lenni Kuff	e97a1b52e0	Remove flaky verification in ALTER/CREATE table tests This fixes the flaky ALTER/CREATE tests by removing a verification step that didn't add value and was non-deterministic. The verficiation step that was removed verified that CREATE/ALTER set the appropriate file format by changing the format to something that didn't match the underlying data files, then attempting to read the data. This is already covered by the positive test case where the file format is changed to match the underlying data. Change-Id: I66f485405234f472f3b83f3e776bf7f2c10de874 Reviewed-on: http://gerrit.ent.cloudera.com:8080/1379 Reviewed-by: Alex Behm <alex.behm@cloudera.com> Tested-by: jenkins Reviewed-on: http://gerrit.ent.cloudera.com:8080/1382 Reviewed-by: Lenni Kuff <lskuff@cloudera.com> Tested-by: Lenni Kuff <lskuff@cloudera.com>	2014-01-28 16:03:02 -08:00
Alex Behm	24662b1941	Allow ALTER TABLE to set per-partition serde and table properties. The main motivation is to allow users to set the per-partition number of rows for manual incremental stats maintenance, as well as a means to 'drop' stats that may have caused undesirable plan changes. Change-Id: Iff38317a993e5d7952ea4df839947f5ec341e930 Reviewed-on: http://gerrit.ent.cloudera.com:8080/1010 Reviewed-by: Alex Behm <alex.behm@cloudera.com> Tested-by: jenkins	2014-01-08 10:54:22 -08:00
Lenni Kuff	39f77b8b8f	Add support for cluster-synchronized catalog operations This change adds support for cluster-synchronized catalog operations. This provides the guaranteethat after a catalog op completes, all other subscribers to the catalog topic have also processed that update. This is useful when load balancing, because a common workflow is to target a different impalad for each statement executed. For example if each of the following were executed sequentially, but targeting a different node: 1) CREATE TABLE Foo 2) INSERT INTO Foo 3) SELECT * FROM Foo 4) INSERT INTO Foo .... Since both the INSERT and the CREATE update the catalog, it would not work as expected without this patch. The user might either get a "table not found" error or would be missing partition information from the INSERT. The downside is that this approach to DDL takes a bit longer because we need to wait until all subscribers have processed an update. If all nodes are healthy, this overhead should not be significantly longer than the current DDL time. However, a single bad node might slow down or completely block the completion of all DDL operations. By default this feature is disabled, but it can be enabled using a new query option: SYNCED_DDL=1 To test this, the base test suite was updated to support selecting a random impalad to execute each query section in a query test file. This is currently only enabled for the insert and DDL tests, but could be leveraged by more tests in the future. TODO: Add additional failure tests around this functionality. TODO: Add an explicit "sync" statement so users do not need to run all their DDL in this mode (since it is slower). Change-Id: I45e757a931bf2a4740cc0cdd1e76ce49a1e22b83 Reviewed-on: http://gerrit.ent.cloudera.com:8080/899 Reviewed-by: Ishaan Joshi <ishaan@cloudera.com> Tested-by: jenkins	2014-01-08 10:53:58 -08:00
Lenni Kuff	8bb0010415	IMPALA-597: Do not crash when multiple partitions have the same LOCATION This patch fixes an issue where Impala would crash if two partitions had the same HDFS location. This is now fixed in hdfs-scan-node. It also includes some cleanup and bug fixes to the FE partition related classes and adds tests. There is still a problem where partition location metadata is not sent to the BE for INSERT statements, but that will be resolved in a separate patch. Change-Id: I0f1c3113d654f7d2b410f00e793ff6b0cae1ae18 Reviewed-on: http://gerrit.ent.cloudera.com:8080/876 Reviewed-by: Alan Choi <alan@cloudera.com> Tested-by: jenkins	2014-01-08 10:53:57 -08:00
Lenni Kuff	e0507e192b	Fix unstable alter table test	2014-01-08 10:50:26 -08:00
Henry Robinson	ead69d377f	IMPALA-249, IMPALA-252: Fixes for static partition keys.	2014-01-08 10:50:14 -08:00
Alex Behm	673d7b97cf	IMPALA-190: Insert with NULL partition keys results in SIGSEGV.	2014-01-08 10:49:22 -08:00
Lenni Kuff	018a72bfe2	IMPALA-189: Properly support NULL partition key values in ALTER .. PARTITION statements	2014-01-08 10:49:21 -08:00
Lenni Kuff	5a0b1270c4	Add support for ALTER ... PARTITION (partitionSpec) SET FILEFORMAT/LOCATION Adds support for: * ALTER TABLE <table> PARTITION (partitionSpec) SET FILEFORMAT * ALTER TABLE <table> PARTITION (partitionSpec) SET LOCATION This enables setting the location and fileformat of specific partitions.	2014-01-08 10:49:17 -08:00
Lenni Kuff	f4a5c0628f	Cleanup HDFS directories before and after running ALTER TABLE tests	2014-01-08 10:49:17 -08:00
Lenni Kuff	1fb72fbc73	IMPALA-156: Support core 'ALTER TABLE' DDL command This patch adds support for - ALTER TABLE ADD\|REPLACE COLUMNS - ALTER TABLE DROP COLUMN - ALTER TABLE ADD/DROP PARTITION - ALTER TABLE SET FILEFORMAT - ALTER TABLE SET LOCATION - ALTER TABLE RENAME	2014-01-08 10:49:14 -08:00

29 Commits