Releases · Unstructured-IO/unstructured

12 Mar 15:57

badGarnet

0.17.0

2dceac3

0.17.0 Latest

Latest

What's Changed

feat: include images when partitioning html by @ryannikolaidis in #3945
fix: pass extract image args to all partitioners by @ryannikolaidis in #3950
feat: allow passing down of ocr agent and table agent by @badGarnet in #3954
Feat/remove reference of PageLayout.elements by @badGarnet in #3943

Full Changelog: 0.16.25...0.17.0

Contributors

badGarnet and ryannikolaidis

Assets 2

07 Mar 11:17

plutasnyy

0.16.25

74b0647

0.16.25

Enhancements

Features

Fixes

Fixes filetype detection for jsons passed as byte streams - Now it prioritizes magic mimetype prediction over file extension when detecting filetypes

Assets 2

07 Mar 11:17

plutasnyy

0.16.24

961c8d5

0.16.24

Enhancements

Support dynamic partitioner file type registration. Use create_file_type to create new file type that can be handled
in unstructured and register_partitioner to enable registering your own partitioner for any file type.
extract_image_block_types now also works for CamelCase elemenet type names. Previously NarrativeText and similar CamelCase element types can't be extracted using the mentioned parameter in partition. Now figures for those elements can be extracted like Image and Table elements
use block matrix to reduce peak memory usage for pdf/image partition.

Features

Add JSON elements to HTML converter - Converts JSON elements file into an HTML file.

Fixes

Assets 2

20 Feb 13:31

plutasnyy

0.16.23

0df50fe

0.16.23

Enhancements

Features

Fixes

Fixes detect_filetype when SpooledTemporaryFile is passed. Previously some random name would get assigned to the file and the function raised error.

Assets 2

20 Feb 01:11

cragwolfe

0.16.22

147add9

0.16.22

Enhancements

Features

Fixes

Fix open CVES in and bump dependencies

Assets 2

17 Feb 16:01

plutasnyy

0.16.21

3403db1

0.16.21

Enhancements

Use password to load PDF with all modes
use vectorized logic to merge inferred and extracted layouts. Using the new LayoutElements data structure and numpy library to refactor the layout merging logic to improve compute performance as well as making logic more clear
Add PDF Miner configuration Now PDF Miner can be configured via pdfminer_line_overlap, pdfminer_word_margin, pdfminer_line_margin and pdfminer_char_margin parameters added to partition method.

Features

Fixes

Fix file type detection for NDJSON files NDJSON files were being detected as JSON due to having the same mime-type.

Assets 2

06 Feb 06:12

cragwolfe

0.16.20

b10379c

0.16.20

Enhancements

Features

Fixes

Fix a security issue where rst and org files could read files in the local filesystem. Certain filetypes could 'include' or 'import' local files into their content, allowing partitioning of arbitrary files from the local filesystem. Partitioning of these files is now sandboxed.

Assets 2

05 Feb 17:21

plutasnyy

0.16.19

5852260

0.16.19

Enhancements

Features

Fixes

Fix a bug where table extraction is skipped when it shouldn't. Pages with just one table as its content or starts with a table misses table extraction. The routing logic is now fixed.
Correct deprecated ruff invocation in make tidy. This will future-proof it or avoid surprises if someone happens to upgrade Ruff.
Remove upper bound constraint on python version in setup.py. Python3.13 is not yet officially supported, but allow users to try.
Fixes removing HTML elements from the inside of table cells in html partition v=2.0. The HTML partitioner now correctly preserves HTML elements from the inside of table cells.

Assets 2

29 Jan 12:52

badGarnet

0.16.17

55debaf

0.16.17

Enhancements

Refactoring the VoyageAI integration to use voyageai package directly, allowing extra features.

Features

Fixes

Fix a bug where build_layout_elements_from_cor_regions incorrectly joins texts in wrong order.

Full Changelog: 0.16.16...0.16.17

Assets 2

27 Jan 23:30

christinestraub

0.16.16

a447b81

0.16.16

Enhancements

Features

Vectorize layout (inferred, extracted, and OCR) data structure Using np.ndarray to store a group of layout elements or text regions instead of using a list of objects. This improves the memory efficiency and compute speed around layout merging and deduplication.

Fixes

Add auto-download for NLTK for Python Enviroment When user import tokenize, It will automatic download nltk data from tokenize.py file. Added AUTO_DOWNLOAD_NLTK flag in tokenize.py to download NLTK_DATA.
Correctly patch pdfminer to avoid PDF repair. The patch applied to pdfminer's parser caused it to occasionally split tokens in content streams, throwing PDFSyntaxError. Repairing these PDFs sometimes failed (since they were not actually invalid) resulting in unnecessary OCR fallback.

Drop usage of ndjson dependency

Assets 2

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

What's Changed

Contributors

0.16.25

Enhancements

Features

Fixes

0.16.24

Enhancements

Features

Fixes

0.16.23

Enhancements

Features

Fixes

0.16.22

Enhancements

Features

Fixes

Enhancements

Features

Fixes

0.16.20

Enhancements

Features

Fixes

Enhancements

Features

Fixes

0.16.17

Enhancements

Features

Fixes

0.16.16

Enhancements

Features

Fixes

Releases: Unstructured-IO/unstructured

0.17.0

What's Changed

Contributors

0.16.25

0.16.25

Enhancements

Features

Fixes

0.16.24

0.16.24

Enhancements

Features

Fixes

0.16.23

0.16.23

Enhancements

Features

Fixes

0.16.22

0.16.22

Enhancements

Features

Fixes

0.16.21

Enhancements

Features

Fixes

0.16.20

0.16.20

Enhancements

Features

Fixes

0.16.19

Enhancements

Features

Fixes

0.16.17

0.16.17

Enhancements

Features

Fixes

0.16.16

0.16.16

Enhancements

Features

Fixes