YAML, SQL, or something else? Looking for recommendations for making a database of stories.

Bubs@lemm.ee · 10 days ago

YAML, SQL, or something else? Looking for recommendations for making a database of stories.

FizzyOrange@programming.dev · 10 days ago

Definitely SQLite. Easily accessible from Python, very fast, universally supported, no complicated setup, and everything is stored in a single file.

It even has a number of good GUI frontends. There’s really no reason to look any further for a project like this.

Bubs@lemm.ee · 10 days ago

One concern I’m seeing from other comments is that I may have more data than SQLite is ideal for. I have thousands of stories (My estimate is between 10 and 40 thousand), and many of the stories can be several pages long.

FizzyOrange@programming.dev · 10 days ago

Ha no. SQLite can easily handle tens of GB of data. It’s not even going to notice a few thousand text files.

The initial import process can be sped up using transactions but as it’s a one-time thing and you have such a small dataset it probably doesn’t matter.

Bubs@lemm.ee · 10 days ago

That’s good to know.

bazzzzzzz@lemm.ee · edit-2 10 days ago

If scraping is reliable, I’d use the classic python pickle or JSON.dump

For a few thousand I would just use a sqlite dB…

3 tables:

Story with fields: Id, title, text
Meta with fields: Id, story-id, subject, contents
Tags with fields Id, story-id, tag

Use SQL joins for sorting etc.

Sqlite is easily converted to other formats if you decide to use more complex solutions.

solrize@lemmy.world · edit-2 10 days ago

Python sqlite3 module for the metadata and it has some features now for full text search that can probably handle a few thousand stories. For a bigger collection like ao3, try solr.apache.org or elastic search etc.

Kissaki@programming.dev · edit-2 9 days ago

I would separate concerns. For the scraping, I would dump data as json onto disk. I would consider the folder structure I put them into, whether as individual files, or a JSON document per line in bigger files for grouping. If the website has good URL structure, the path could be useful for speaking author and or id identifiers in folders or files.

Storing json as text is simple. Depending on the amount, storing plain text is wasteful, and simple text compression could significantly reduce storage size. For text-only stories it’s unlikely to become significant though, and not compressing makes the scraping process, and potentially validating completeness of scraped data simpler.

I would then keep this data separate from any modifications or prototyping I would do regarding modification or extension of data and presentation/interfacing.

Bubs@lemm.ee · 8 days ago

After reading some of the other comments, I’m definitely going to separate the systems. I’ll use something like json or yaml as the output for the raw scraped data, and some sort of database for the final program.