W3C

Publishing Maintenance Working Group Telco

10 September 2026

Attendees

Present
Avneesh Singh, Charles LaPierre, Dale Rogers, Brady Duga, Gautier Chomel, George Kerscher, Gregorio Pellegrino, Hadrien Gardeur, Ivan Herman, Daniel Kimberg, Laurent Le Meur, Masakazu Kitahara, Shinya Takami, Susan Neuhaus, Toshiaki Koike, Wendy Reid
Regrets
-
Chair
Susan Neuhaus, Wendy Reid
Scribe
Susan Neuhaus

Meeting minutes

Annotations

Laurent Le Meur: This is where we left off for the summer, but there are some issues we can tackle…
… I'm still looking for implementers, Thorium will get annotations in 3.6 in 2 months as a test
… I hope someone from Readwise or colibrio will do also

w3c/epub-specs#2852

Laurent Le Meur: This issue was offering breadcrumbs to the reading system to get back to the origin

<Ivan Herman> +1

Laurent Le Meur: since no one expressed a need for this I propose we remove issue 2852

Hadren: we have the body, and this would block other contextual elements, there wouldn't be any human readable code left if we drop this

Laurent Le Meur: If you drop the HTML you would still have something you could export from the locator, like the publisher and the book

Hadrien Gardeur: I am not a fan of having something specific, a general text can be most useful. It could be a title, something from the TOC
… what I'm describing is a little different than what you have here. You could resolve it is a few ways
… when you have access to the file it is possible to extract the information, but when you have only the extraction, it could be come a problem. We could use something broader in scope

Ivan Herman: are we talking about the same thing? On the annotation set there is an about object, and the issue isn't about this. The meta object was created on the target which means something undefined

Hadrien Gardeur: I was talking about an annotation that is not on the annotation set level
… basically I'd like to make sure annotations are useful even when I don't have access to the original publication

Laurent Le Meur: If we want that it means replacing the structure with a string giving context, like chapter title or page. It would be up to the reading system to decide what that is

Brady Duga: I agree with Hadrien Gardeur, but is this something I need to do for version 1 of the spec. It feels useful, but it will take some work to figure out the right way to do it.
… if we're not specific, this could be too obscure to be used. If we don't want to do this for 1.0 we should remove it

<Dale Rogers> s/same thing./same thing?

Hadrien Gardeur: we work with the publication timeline, like at all times I have a header or a progress bar, so I know where I am in a publication. Also useful in a search
… a publication timeline is easy to generate with prepaginated media, and audio or video. Its more complex with a reflowable ebook but it is useful. And helpful for the user.

<Dale Rogers> s/same thing./same thing?/

Hadrien Gardeur: this string would be similar to what we would generate for a timeline. Each reading system does it their own way

Ivan Herman: to be practical. even though we haven't talked about a timeline, I don't think this specification will make it to recommendation by the end of this working group (February '26). Se we can extend the timeline and come up with a proper set of attributes.
… we should acknowledge that this is open and needs more work. Someone could take on finding more specific properties.

Laurent Le Meur: I am OK with this proposal

w3c/epub-specs#2884

Laurent Le MeurLM: The target may not be a specific segment in the document unless we consider that the whole document can be a specific segment
… the problem comes from the word "segment" and what can be the size of the target. We know if can be a video or audio. The term segment comes from the W3C model itself.
… I proposed a change of working to "segment of interest" I'm looking for ideas to rewrite that

Hadrien Gardeur: by segment do you mean fragment?

Laurent Le MeurLM: perhaps we should remove the sentence, we should just say what it is. There are restraints in our annotation beyond what is in the W3C annotations

Brady Duga: There are nuances here, and we can delete this sentence. We just have to explain how things are different.

w3c/epub-specs#3009

Ivan Herman: This is a general thing, what should be listed as metadata for an annotation set. What we have now is not much

Laurent Le MeurLM: you prefer we remove "generator" it was part of the W3C model, I see it as a useful tool
… you also propose to remove the DC format, because we shouldn't speak about format in this specification
… the properties retained are identifier, format, title, publisher, creator, date of release. These are common properties to find in epub
… you propose to remove format, why?

Ivan Herman: my question about generator, is to define what it should be, I don't know how I would debug this and what I would see.

Laurent Le MeurLM: I see, I agree

Brady Duga: I am against putting debugging information in any export format, it doesn't belong in the spec it is a privacy issue

Ivan Herman: that is in line with what I'm saying, put in a URL.

Brady Duga: If I use a screen reader, I may not export that information to everyone

Ivan Herman: about the format issue, it is a question we should answer ourselves
… I have no idea how PDF works, if all the things we define are usable in PDF and we are not able to confirm this. PDF is irrelevant unless we go through the whole document and add PDF specifications
… that's why format isn't useful, because the only value would be EPUB

Gautier Chomel: I might need to know the format of the output for use case 5,3 Annotations used in the publishing workflow

Hadrien Gardeur: I don't think we can trust all the metadata that we have, we probably need less. I'm OK with title and time stamp. DC date seems weird, we may not have this information in a reliable place.

Wendy Reid: I agree with Hadrien Gardeur, there are contexts inwhich ID is important, these may be context specific, this information may be less reliable outside of that rs
… I see what you are saying Gautier Chomel, I think some of what you mention could be inferred from the file itself, and other information might escape

Ivan Herman: we are talking about all this in isolation, when I export information, these metadata values will be in the exported package, we need to say that these must be a copy of the values use in the document
… and this is testable for conformance
… the question is whether the importing system can trust it, and maybe not, but it provides information

Laurent Le MeurLM: I agree there is an advantage to exporting information about the document in the annotation set
… when I am importing annotations from one system into the same title on another system, the information would be helpful. Especially if a human can make the decsion

Hadrien Gardeur: we know unique identifiers are not always unique. I think this will fail. I expect the real use case will be that people will import annotations for a particular book. I don't expect annotations to translate well from one system to another.

Brady Duga: I agree with Hadrien Gardeur about the identifier, and also that there might be information useful to the user, we may want to be able to give the user information about the annotations before they import them.
… some of the information is useful, but DC identifier is useless

Ivan Herman: what if I take an identifier like a hash of the whole package document, then I can check if the origin of the annotations and the current document are the same

Brady Duga: I don't know how well that will work in practice

Wendy Reid: I agree knowing dc identifier is unreliable, the uuid can be too specific, but that can be updated per edition, what if we identify the publication through multiple parts of the metadata in a decsending order.
… ultimately we have something that is useful to show the reader before they import it. We can use title, publisher, and more information from the reading system like related titles.

Hadrien Gardeur: We've seen examples of people trying to use hash and it fails. One example is KO reader sync, they send a hash and a path or two paths. Whenever you do anything with the file the hash changes.
… hash is a brittle thing to use, like a house of cards. I like Brady Dugas suggestion of using the minimum information

George Kerscher: are we going to have enough information to create a bibliographic reference?
… there is more information needed than the dc:title

Wendy Reid: we are limited to what's in the package document, not all publishers put in good data

Hadrien Gardeur: if we are talking about citation references, there are many styles, and it is unlikely that we would get all that data in an epub file.

Laurent Le MeurLM: if we take the baseline, then we include only title and creator in the metadata

Wendy Reid: maybe publisher

Laurent Le MeurLM: publisher doesn't seem useful to the user. Even if there are two books with the same title the creator will help differentiate

<Susan Neuhaus> +1 Wendy Reid

Dale Rogers: from an annotation point of view, knowing which book information to put into an annotation set, does a reading system have to check the incoming information against an existing title. Is the issue at the software level or the user's level?

Laurent Le MeurLM: a good use case: I am a user and I want to associate annotations with a certain book. I select the epub, select the annotations, then the reading system checks the title, and if there is a discrepancy, it alerts the reader. A human choice aided by the machine

Brady Duga: I don't know enough about dc:publisher in ebooks if it is useful or not. As a reader I don't care who published my book. But it might be important to other readers. Since this isn't usually exposed to the user, and if the data isn't good, it could be confusing. I lean toward not including it but could be talked out of it

Ivan Herman: I am worried about the way we go into this discussion. We want to be sure none of the information can be misinterpreted. But then it doesn't get used because we are worried. In a large number of cases, if I know author, publication, year, etc. It will give me enough information to find it. We are throwing away everything because there are some cases where it doesn't work.
… I think we should make it clear to publishers that the package document information should be there, even if in practice these things can go wrong. Let's not throw away the information because sometimes it can go wrong.

Gautier Chomel: we should encourage good practices in publishing and not punishing the ones who do it right

Laurent Le MeurLM: I made a mistake in the dc:date, it should be the date the Ebook is published

Dale Rogers: for an ebook we would tell our students where to get the book. I know everyone would get the same epub, same date, etc. So outputing the annotations in a classroom situation would work well. But could get more complicated in other use cases.

George Kerscher: students will many times not like the school reading system environment and get the book from bookshare or another place and use the reading system that works for them. Happily if the publisher provided the title to bookshare the metadata should be the same.

AOB

<Wendy Reid> https://docs.google.com/document/d/1FfUjiK8PrKfqeVCAnjpgZ_rpq7ZSAziloyLT5AkZKjQ/edit?tab=t.0#heading=h.v52yw3mb2h7

Wendy Reid: we will talk about this more, here is a link to the TPAC agenda for you to review.

Minutes manually created (not a transcript), formatted by scribe.perl version 244 (Thu Feb 27 01:23:09 2025 UTC).