109
edits
Changes
From Final Fantasy Inside
m
This section explains how LGP is an archive format used by Final Fantasy VII to store game assets. The format is used in the LGP archives from FF7 PC are constructed. If you're looking for a tool that already manages LGP archivesversion of the game to package various game resources including textures, models, scripts, try [[User:Ficedula|Ficedula]]'s [http://sylphds.net/f2k3/index.html LGP Editor]and other data files.
Essentially the An LGP file is split up into four archive consists of a header, table of contents (maybe lessTOC), hash lookup table, optional path table for directories, depending on how you count it) sectionsand the actual file data. All multi-byte integers in the format are stored in '''little-endian''' format.
# File header/Table The archive terminates with the magic string "FINAL FANTASY7" to mark the end of contents# CRC code# Actual data# File terminatorthe file.
This contains two partsAn LGP archive is divided into six sections in the following order: A header of fixed size, then the table of contents.
Next is a four'''Notes:'''* Filenames are stored without directory paths* Filenames are null-terminated within the 20-byte integer saying how many files field* The offset field points to the start of the File Header for this file* The path field references the Path Table (1-indexed); 0 means the archive contains.file is in the root directory* Maximum filename length is 19 characters plus null terminator
Following this is the table of contents (TOC): One entry per file.== Hash Table ==
This section # Compute hash from the filename (first two characters of stem)# Read HashTable[hash]# If index is 3600 bytes. It is 30 sets of 30 0, file not found (empty bucket)# Search TOC entries containing two 16starting at index-bit words each 1 (30 x 30 x 2 x 2 = 3600converting to 0-indexed)for count entries# Match by exact filename comparison
The sets contain file'''Example:''' To find "test.dat":# Compute hash: hash('t') × 30 + hash('e') + 1 = 19 × 30 + 4 + 1 = 575# Read HashTable[575]: {index: 42, count: 3}# Search TOC entries 41-43 (0-indexed: index-group information which is based on the first two letters of each file 1)# Find entry where name== "test.dat"
The first letter, minus the value for ascii 'a' (0x61) is the index of the set to which the file belongs.== Path Table ==
Each entry is two words.=== Path Table Header ===
The first word is the 1{| class="wikitable"! Offset !! Size !! Type !! Description|-based index | 0x00 || 2 bytes || uint16 || Number of the directory entry for the first file in the set.path groups|}
The second word is the number of files in the set, most of which are 0x003c (60). There are a few entries after the bulk which have fewer entries.=== Path Group ===
The meaning of these sets and why they're divided in this manner is yet to be determined.Each path group contains all directory paths for files sharing the same filename:
There is one 16{| class="wikitable"! Offset !! Size !! Type !! Description|-bit word with the value | 0x00 || 2 bytes || uint16 || Number of 0 (0x0000) at the end paths in this group|-| 0x02 || Variable || Path[] || Array of this data which may belong to this section or the next.Path entries|}
The data from the files. However it's not that simple: the TOC doesn't list how long each file is (somewhat useful). It's done here. The offset in the TOC is actually the position of yet another file header. Format Each path entry is130 bytes:
After the last piece To read an LGP archive: # '''Read Header''' - Validate magic string "SQUARESOFT" and extract file count# '''Read TOC''' - Read file_count entries of 27 bytes each, storing filename, offset, type, and path index# '''Skip Hash Table''' - Skip 3600 bytes (or read for validation/fast lookup)# '''Read Path Table''' - Read path group count, then for each group read path count and path entries# '''Resolve Full Paths''' - For each TOC entry with path index > 0, look up path group at path_index-1, find path entry matching TOC index, and prepend path to filename# '''Read File Data''' (on demand) - Seek to TOC entry's offset, read 24-byte file header, read file_size bytes of data comes the content == Writing Algorithm == To create an LGP archive: # '''Build Metadata''' - Compute hash for each file descriptor. This is a simple , group files by hash bucket, identify files needing path entries# '''Write Header''' - Write magic string, except instead of being nullfile count, and reserved fields# '''Write TOC''' - Write entries with placeholder offsets (will be updated later)# '''Write Hash Table''' -terminated itPopulate 900 entries based on file hash groupings# 's terminated ''Write Path Table''' - Group paths by the end of the filename collisions and write path groups# '''Write File Data''' - For each file. Itwrite header and content, recording actual offset# '''Update TOC''' - Seek back to TOC section and update offsets with actual values# '''Write Footer'''s - Append "FINAL FANTASY 7FANTASY7" for all archives, except terminator == Size Limits == The LGP patches, where it's format has the following technical limitations: {| class="LGP PATCH FILEwikitable".! Limit !! Maximum Value !! Reason|-| Files per archive || 65,535 || uint16 in header|-| File size || 4 GB || uint32 in file header|-| Archive size || 4 GB || uint32 offsets in TOC|-| Filename length || 19 characters || 20 bytes with null terminator|-| Path length || 127 characters || 128 bytes with null terminator|-| Hash buckets || 900 || Fixed 30 × 30 grid (letters × letters)|}
==== Notes ====
→Introduction
=== LGP Archive format for PC by [[User:Ficedula|Ficedula]] =Introduction ==
==== Section 1: File Header ==Structure Overview ==
{| class="wikitable"! Section !! Size !! Description|-| Header || 16 bytes || Archive metadata and file count|-| Table of Contents || 27 bytes × file count || File entries with names and offsets|-| Hash Table || 3600 bytes (fixed) || Lookup table for fast file access|-| Path Table || Variable || Optional directory paths for files|-| File Data || Variable || Actual file contents with individual headers|-| Footer || 14 bytes || "FINAL FANTASY7" terminator string|} == Header == The first item header is 12 always 16 bytes containing and contains the archive's magic identifier and file creatorcount. This is a standard {| class="wikitable"! Offset !! Size !! Type !! Description|-| 0x00 || 2 bytes || uint16 || Reserved (always 0)|-| 0x02 || 10 bytes || char[10] || Magic string, except it is : "rightalignedSQUARESOFT". In other words the blank space comes before the actual text(ASCII, not after. In FF7 itno null terminator)|-| 0x0C || 2 bytes || uint16 || Number of files in archive|-| 0x0E || 2 bytes || uint16 || Reserved (always 0)|} '''Validation:'''s always The magic string at offset 0x02 must exactly match "SQUARESOFT" preceded by two nulls to make it 12 (ASCII, no null terminator within the 10 bytes). The only other thing you might see is == Table of Contents (TOC) == Immediately follows the header and contains one 27-byte entry for each file in the archive. {| class="wikitable"FICEDULA! Offset !! Size !! Type !! Description|-| 0x00 || 20 bytes || char[20] || Filename (null-LGP"padded, which I use to indicate a no path)|-| 0x14 || 4 bytes || uint32 || Absolute file is an LGP *patch* one of my programs has constructedoffset in archive|-| 0x18 || 1 byte || uint8 || File type (always 0x0E / 14)|-| 0x19 || 2 bytes || uint16 || Path index (0 = no path, not a complete archive.1+ = path table index)|}
The hash table immediately follows the TOC and is always exactly 3600 bytes (900 entries × 4 bytes). It provides fast O(1) lookup of files by filename. === Hash Table Entry === Each entry in the TOC has the following structureis 4 bytes:
{| class="wikitable"
! Offset! Length! Size !! Type !! Description
|-
| 20 0x00 || 2 bytes| Null terminated string| uint16 || Index into TOC (1-indexed, giving filename0 = empty bucket)
|-
| 4 byte integer0x02 || 2 bytes || uint16 || Position Count of consecutive entries in this bucket|} === Hash Function === The hash is computed from the filename (without path or extension) using only the first two characters of the file where data starts for stem: <pre>hash_value = hash(first_char) × 30 + hash(second_char) + 1</pre> For filenames with only one character in the filestem: <pre>hash_value = hash(first_char) × 30</pre> === Character Hash Values === The hash function maps characters to numeric values as follows: {| class="wikitable"! Character !! Hash Value
|-
| 1 byte| style="background: rgba-z (255,255,204case insensitive)" | Some sort of check code. File attributes? Normally seems to be<br />14 but it does vary.| 0-25
|-
| 2 byte short0-9 || 0-9|-| _ (underscore) || 10 (same as 'k')| style="background: rgb-| - (255,255,204hyphen)" | Something to do with duplicate file names. If a name is unique it is 0, otherwise it is assigned a value based on existing duplicates. | 11 (Hard to explainsame as 'l')
|}
'''Note:''' The hash function is case-insensitive, treating 'A' and 'a' identically. ==== Section 2: Section formerly designated as "CRC Code" =Lookup Algorithm ===
The second letter, minus path table immediately follows the value hash table and stores directory paths for ascii ' ' ' (0x60) is the index of the entry within the setfiles. Since the second letter in all of the file names is ascii 'It uses a' or greater, it means variable-length structure organized into "path groups" for files that share the first entry same filename but exist in every group is always zero (0x0000)different directories.
==== Section 3: Actual Data =Path Entry ===
{| class="wikitable"
! style="background: rgb(204,204,204); width: 80px" align="center" | Offset !! Size! style="background: rgb(204,204,204); width: 200px" | ! Type !! Description
|-
| 0x00 || 128 bytes || char[128] || Directory path (null-padded, no trailing slash)|-| 0x80 || 2 bytes || uint16 || TOC index this path belongs to|} '''Notes:'''* Path index in TOC entries is 1-indexed into path groups* Multiple files with the same name but different paths share a path group* Empty path string means root directory* Maximum path length is 127 characters plus null terminator == File Data == File data blocks follow the path table. Each file consists of a 24-byte header followed by the raw file content. === File Header === {| class="wikitable"! Offset !! Size !! Type !! Description|-| 0x00 || 20 bytes|| char[20] || Filename (same as in TOC)| Null terminated string, giving filename-| 0x14 || 4 bytes || uint32 || File size in bytes|} === File Content === Immediately follows the file header. The size is specified in the header's file size field. {| class="wikitable"! Offset !! Size !! Type !! Description
|-
| 4 bytes0x00 || file_size || byte[] || Raw file data|} == Footer == The archive ends with a 14-byte terminator string: {| File lengthclass="wikitable"! Offset !! Size !! Type !! Description
|-
| Varies0x00 | The file data itself| 14 bytes || char[14] || "FINAL FANTASY7" (no null terminator)
|}
==== Section 4: Terminator ==Reading Algorithm ==
The game is remarkably flexible about LGP archives. So long as the TOC and the CRC data is intact it'll accept just about anything.