# How this site was reconstructed

*Plain-text version of the disclosure published at `/about/` in the reconstruction.
Figures are read from the project ledger, not transcribed by hand.*

> **This document describes the reproduction, not the original site.** Everything
> described here — the banner at the top of each page, the About page, the snapshots
> behind outbound links — belongs to the 2026 reconstruction. The original site,
> published in September 2009, contained none of these elements.

*La Universidad Hoy* was the public site of the University of Puerto Rico presidential
transition, published in September 2009 at `hoy.upr.edu`. The project was guided by **José Luis
Cruz Rivera**, who led the presidential transition team while serving as Vice President
of Student Affairs and coordinator of the university's strategic plan, *Diez para la
Década*, until the end of Antonio García Padilla's presidency on September 30, 2009.

Neither the domain nor the platform survives: `hoy.upr.edu` no longer resolves, and
classic Google Sites — the platform it was built on — was retired in 2021. What is
reproduced here comes from captures preserved by the **Internet Archive's Wayback
Machine**.

---

## 1. Reference

| | |
|---|---|
| Original site | `hoy.upr.edu` — *La Universidad Hoy*, published September 2009 |
| Native address | `sites.google.com/a/upr.edu/launiversidadhoy` |
| Institution | University of Puerto Rico |
| Source of this reproduction | Wayback Machine, Internet Archive (`web.archive.org`) |
| Capture range | 2009-09-29 — 2024-08-05 |
| Date of reconstruction | 2026-07-28 / 29 |

## 2. Authorship and working method

The reconstruction was carried out in a single working session between **José Luis Cruz
Rivera** and **Claude**, Anthropic's coding assistant, using Claude Code.

| | |
|---|---|
| Planning and method design | Claude Fable 5 |
| Implementation and execution | Claude Opus 5 |
| Direction, decisions and validation | José Luis Cruz Rivera |
| Date | 2026-07-28 / 29 |
| Session | 2026-07-28 21:15 — 2026-07-29 06:45, MST (UTC−7), Flagstaff, Arizona |
| Exchanges in the conversation | 24 |

The method was designed and reviewed before any code was written. The entire process —
harvesting, verification, assembly and reporting — is automated in programs that can be
re-run; no file was edited by hand.

## 3. Sources

The site was reachable at two equivalent addresses. Captures were indexed from every
address it ever used.

| Source | Archived period | Records |
|---|---|---:|
| `hoy.upr.edu` (custom domain) | 2009-09-29 → 2014 | 517 |
| `sites.google.com/a/upr.edu/launiversidadhoy` | mostly the November 2020 sweep before Google Sites closed | 2,643 |
| `*-sites.googlegroups.com` (attachment hosts) | reached by following archived redirects | — |

The earliest capture of any kind is **September 29, 2009**. Nothing before that date
survives in any archive, so the site's first months are beyond the reach of this or any
other reconstruction.

## 4. Fixity verification

The Internet Archive's CDX index publishes its own SHA-1 checksum for each capture. That
checksum was recomputed on every downloaded file and compared against the published one.

| | |
|---|---:|
| Successful downloads | 1,422 |
| SHA-1 recomputed and matching the Internet Archive's own | 911 |
| HTML pages downloaded | 684 |
| HTML pages checksum-verified | 682 (99.7%) |

PDF documents do not admit this check, for a technical reason: classic Google Sites never
served an attachment directly. Every document link was an HTTP 302 redirect into a
signed, expiring address on a `googlegroups.com` host. The checksum the index publishes
for such a record describes **the redirect, not the document**. Those files were validated
structurally instead: binary signature (`%PDF`), content type and length.

Two late HTML captures did not match their published checksum. Both are named here so the
claim can be checked:

| page | capture | archive's SHA-1 | recomputed |
|---|---|---|---|
| `proyecciones-sistemicas/oficina-de-desarrollo-fisico-e-infraestructura/recinto-de-ciencias-medicos` | 2022-08-31 | `4DNKQD4FKQH6…` | `7DRAK6L5J2YN…` |
| `proyecciones-sistemicas/oficina-de-desarrollo-fisico-e-infraestructura/upr-cayey` | 2023-06-05 | `A3RVNWAFCYVC…` | `LH3VLZTOTL57…` |

**Neither is used in the reproduction.** The first page is drawn from a November 2020
capture and the second from a January 2010 one, both of which verified. Every page shown
comes from a checksum-verified capture.

## 5. Capture selection

The site was published once, in September 2009, and not revised afterwards. That account was tested rather
than assumed. Raw HTML always differs across years because Google rewrote its own page
chrome, so the comparison was made on **extracted text**, not markup.

Of 65 pages captured more than once, **two content pages differ**, and the difference runs
in the direction of loss: the September 29, 2009 capture carries navigation entries that
later captures no longer have — *Blog*, *Diez para la Década*, *Avanza*, and *La
Universidad: Inversión Estratégica de Puerto Rico*. The other four differing pages are Google's own dynamically
generated pages (recent changes, search, sitemap), not site content.

**Decision:** the earliest verified capture of each page is used, being demonstrably the
fullest and closest to publication.

**The consequence, stated plainly:** the September 2009 home page did not yet link the
`[2008-2009]` annual report, which appears in captures from January 2010. Since the
document itself was recovered, leaving it unlinked would have made it present but
unreachable, so that one link was added and is disclosed as modification 9 below. No
other later-added material was carried back. The `versions/` set preserves **every**
distinct capture of every page, unmodified.

Because the early crawler did not reach every page, the reproduction does not come
uniformly from 2009. By capture year: 2009 (6 pages), 2010 (36), 2011 (10), 2012 (8),
2019 (1), 2020 (35). Since those pages' text does not differ across captures, this affects
the recency of the capture, not the content.

## 6. Register of modifications

The complete list of differences between the original files and what is shown.

### Structural — required for the site to work offline

1. **Link rewriting.** Absolute addresses pointing at `hoy.upr.edu` or the native address
   were converted to relative paths. Link text, order and destination are unchanged.
2. **Character-set declaration.** The archive served `charset=utf-8` in the HTTP header
   only; the pages carry no `<meta charset>`. A static file server does not reproduce that
   header, so accents rendered as corrupted text (`EstratÃ©gico`). A single
   `<meta charset="utf-8">` tag was inserted into each page. **This is a tag that was not
   in the original bytes**; it affects encoding interpretation only, and alters no
   character of the text.
3. **Theme resources.** Stylesheets, scripts and theme images were served from
   `www.gstatic.com`. They were downloaded **from the Internet Archive** (not from Google
   today) and are served locally — 909 such assets.

### Content — declared divergences

4. **One video restored from a personal copy.** The "Mensaje del Presidente" panel was a
   Google Sites gadget embedding a YouTube video (id `VDtJX5AQ9Iw`). It was never a file on
   the site, so no crawl could recover it; the video has since been deleted from YouTube
   and was never archived there either. José Luis Cruz Rivera supplied the original MP4
   from a personal backup. It is shown in its place, marked in the markup as
   `data-provenance="personal-backup"` and recorded separately: **it is not counted as
   material recovered from the archive**. It is the only file in the reproduction that did
   not come from the Internet Archive.
5. **One broken link deactivated.** The home page linked `UPRINVERSIONESTRATEGICA.pdf`
   ("La Universidad: Inversión Estratégica de Puerto Rico"). That document exists in no
   archive under any address form. Its anchor was converted to plain text: **the text
   remains visible on the page**; only the unusable hyperlink was removed. The original
   markup is preserved unmodified, and the document remains listed in the gap report.
6. **Two links marked unavailable.** `junta-de-subastas` and `junta-de-apelaciones-1`
   appeared in the site's navigation from 2009, but **no crawler ever captured either
   address**. When one was first attempted — July 2011 for `junta-de-apelaciones-1`,
   November 2020 for `junta-de-subastas` — it returned HTTP 404, so by then the page was
   gone. Neither appears in the sitemap captures from October 2009 onward. **Whether
   either page existed earlier and was later removed cannot be determined from the
   archive**: the gap between the link appearing and the first crawl attempt is two years
   in one case and eleven in the other. They open a page stating exactly that. An earlier
   version of this reconstruction pointed them at each body's present-day page on
   `upr.edu`; that was withdrawn, because offering a 2025 page in place of a lost one
   papers over the loss even when labelled.
9. **The 2008-2009 annual report, re-linked.** This document was not available when the
   site was published. José Luis Cruz Rivera submitted it for publication later, once it
   had become available and he had left his position, which is why the link appears in
   captures from January 2010 onward and not on the September 2009 home page this
   reconstruction is built from. The document itself was recovered, so leaving it unlinked
   would have made it present but unreachable from the navigation. It was added to the
   "Informes Anuales UPR" list, carrying `data-added="1"` and a tooltip giving the reason.
   **This is a link the September 2009 page did not have.** No other later material was
   carried back.
7. **Outbound links open period captures.** The site linked to eight UPR properties in its
   sidebar and, throughout its pages, to campuses, accreditors and services. Those hosts
   still exist but show 2026 content. Each outbound link opens a local page showing that
   site as captured closest to September 29, 2009, **with the capture date stated in the
   caption**, plus a link to the live site.

   **Only captures from the era are shown** — 2012 or earlier. A later capture is not a
   record of what the site linked to, so where the archive holds nothing from the period
   the page says so rather than showing a 2013 or 2025 version; Radio Universidad
   (`wrtu.pr`, earliest capture 2013) is the one destination this affects. 58 snapshots
   are displayed, all of them from the era. Four predate the site's publication, which
   the caption notes. Every screenshot carries a standing note that an archived capture
   shows only what the crawler saved, since images the archive never fetched cannot be
   recovered afterwards.
8. **Reconstruction notice and the About page.** A banner was added at the top of every
   page, deliberately outside the site's own visual language, identifying the document as a
   reproduction and linking to the disclosure. Neither the banner nor the About page was
   part of the original site.

### What was not done

No text was edited, rewritten, corrected or reordered. No document was altered: every PDF
is byte-identical to what the archive served. No dates, authorship or metadata were
changed. No new content page was written, and no content was added beyond items 4 and 9. Nothing
was removed from the record: the text of the single deactivated link remains visible, and
every gap is enumerated in the accompanying report.

## 7. What is missing

**710 of 729 resources (97.4%)** of everything the site offered were recovered: 483 of
485 documents the site hosted itself, and 82 of the 83 pages listed in its own sitemap
(the remaining one is a blank Google Sites template).

Measured only over what the site hosted, the figure would be **710 of 713 (99.6%)**. That
is not used as the headline because it leaves out 16 documents the site published from a
Dropbox public folder, which are equally gone. To a reader clicking the link, the document
fails to appear either way, so they are counted as missing.

Not recovered:

0. **One recovered from a third source.** The 2008-2009 Senate memorial was
   found in the Puerto Rico Office of Management and Budget's virtual library and
   is included, labelled `provenance=external-source` — neither archive material
   nor personal backup. `reports/SEARCH_RECORD.md` documents the whole search,
   including the avenues that came up empty, so the effort is not repeated.
1. **15 documents hosted outside the site** — the rest of the *Memorial del Presidente de la UPR —
   Presupuesto* series (7 Cámara + 7 Senado), plus `3DTimeline.mov`. They
   sat in a Dropbox public folder; Dropbox retired those links in 2017. The CDX index was
   queried for each address, encoded and unencoded: **all 16 return zero captures**, so
   the archive never held them — they were direct downloads behind a redirect, which
   crawlers rarely follow. That result is recorded per-destination in the ledger
   (`outbound.archived_captures`) rather than asserted. Full list in
   `reports/MISSING_DOCUMENTS.md`.
2. **`UPRINVERSIONESTRATEGICA.pdf`** — never captured by any crawler. It was linked once,
   from the home page, only in the September 29, 2009 capture.
3. **Two navigation targets** never captured by any crawl, and already returning 404 by
   the time a crawler first tried them (item 6.6).
4. **The embedded video**, recovered from a personal copy rather than an archive (item 6.4).

Anything published and withdrawn before September 29, 2009 is unrecoverable, since no
capture predates that date — but since the site went up that same month, this window is
days wide, not months.

## 8. How to verify

**Check any document.** `inventory.csv` lists every recovered file; the project ledger
records, for each, the exact Internet Archive address and capture date it came from. That
address can be fetched today and compared byte for byte.

**Check the page list.** The site's own sitemap is archived, and its embedded code
enumerates every page it contained. That — not an inference from links — is the
denominator the completeness figures are measured against:

```
https://web.archive.org/web/20091003125041/http://hoy.upr.edu/system/app/pages/sitemap/hierarchy
```

**Check that a document is missing rather than omitted.** The Internet Archive's CDX index
can be queried directly; it returns nothing for the absent document under any address form:

```
curl "http://web.archive.org/cdx/search/cdx?url=upr.edu&matchType=domain&filter=urlkey:.*inversionestrategica.*"
```

**Rebuild from scratch.** `./run_pipeline.sh` re-derives everything from the Wayback
Machine. The reproduction is a *derived* artifact; `raw/blobs/` and `versions/` hold the
unmodified source bytes.

## 9. Evidence bundle

| file | contents |
|---|---|
| `inventory.csv` | Complete inventory of recovered files |
| `gaps.csv` | Items not recovered, with classification |
| `link_audit.csv` | Audit of every link in the reproduction |
| `snapshots.csv` | Period snapshots of outbound links |
| `CHANGE_CHECK.md` | Text comparison across captures |
| `GAPS.md` | Gap report |
| `provenance.jsonld` | Machine-readable provenance (W3C PROV-O) |
| `launiversidadhoy-evidence-bag.zip` | The whole package as a conforming BagIt bag (RFC 8493) |
| `manifest-sha256.txt` | SHA-256 checksums of the loose files |

## 10. Standards followed

The structure of this disclosure follows established digital-preservation practice, so it
can be assessed against recognized criteria.

| standard | role here |
|---|---|
| OAIS — ISO 14721 | Organization of preservation description information: reference, context, provenance, fixity, access |
| PREMIS 3.0 | Preservation metadata vocabulary: agents, events, fixity checking via cryptographic checksums |
| W3C PROV-O | Machine-readable provenance (entities, activities, agents), included as JSON-LD |
| BagIt (RFC 8493) | The evidence package is shipped as a conforming bag: `bagit.txt`, `bag-info.txt`, payload manifest, tag manifest, payload under `data/`. Validatable with `bagit.py --validate` |
| Memento — RFC 7089 | Citation of archived resources by capture datetime |

## 11. Known limitations

**The video rests on testimony, not on an archive.** It is the only object here without
externally checkable provenance. The original was on YouTube, has been deleted, and the
Internet Archive never captured it; the only trace in the archived pages is the video id,
with no duration or other metadata to compare against. There is no independent way to verify
that this file is that video. Its fixity is recorded so it cannot change from here on:
27,722,817 bytes, MP4 v2 (ISO/IEC 14496-14), SHA-256
`79de0fe605248c01eca3c662b77d8ce34f4fa0b315fe13aedf9590ec4535adca`.

**No original content passed through a language model.** Saying no file was edited by hand
does not cover this. The models wrote the programs, the About page and the interstitials.
No byte of the original site's content passed through a model: the programs copy bytes from
the archive to disk and verify them, and content is never read into a model, summarised,
regenerated or corrected.

**A single session, with no independent review.** Eleven continuous hours, no peer review.
The process is reproducible end to end and every claim can be checked against source, but
that is not a second pair of eyes. Such review is invited.

**The denominator rests largely on one source.** The page list comes from the site's own
map, cross-checked against the link graph from every capture. The two agree, but the map is
the authority and survives in a single capture. An independent inventory, if one surfaced,
would be worth reconciling.

**The reconstruction can be lost too.** It is now archived in the Wayback Machine
(2026-07-29). Depositing the code and the evidence bag in repositories with permanent
identifiers remains outstanding.
