A small regression test for Chinese and Japanese document ingestion: catch whitespace-based chunking, add a code-point cap, and preserve source offsets without confusing characters with embedding tokens.
A small regression test for Chinese and Japanese document ingestion: catch whitespace-based chunking, add a code-point cap, and preserve source offsets without confusing characters with embedding tokens.