<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.1">Jekyll</generator><link href="http://localhost:4000/feed.xml" rel="self" type="application/atom+xml" /><link href="http://localhost:4000/" rel="alternate" type="text/html" /><updated>2026-03-07T02:00:07+01:00</updated><id>http://localhost:4000/feed.xml</id><title type="html">The Fox &amp;amp; Badger</title><subtitle>Personal professional blog of dr. Joris J. van Zundert, senior researcher at the Huygens Institute in Amsterdam. Opinions and content do no reflect Huygens Institute policies or consent. All content CC-BY unless stated otherwise.</subtitle><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><entry><title type="html">Too Good to be Accepted</title><link href="http://localhost:4000/to-good-to-be-accepted/" rel="alternate" type="text/html" title="Too Good to be Accepted" /><published>2026-03-04T19:12:25+01:00</published><updated>2026-03-04T19:12:25+01:00</updated><id>http://localhost:4000/to-good-to-be-accepted</id><content type="html" xml:base="http://localhost:4000/to-good-to-be-accepted/"><![CDATA[<p>TL;DR You have been rejected at <a href="https://dh2026.adho.org/">DH</a>. This may well be because you are more expert than many DH peers and your work is more advanced in a very particular niche. There is no hard data on this type of rejection, but it would be interesting to figure out how to do the bibliometrics and actually do them. Anyway, do not sulk, do not mock. Submit to satellite conferences such as <a href="https://2027.computational-humanities-research.org/">CHR</a> and <a href="https://jcls.io/site/ccls2026/">CCLS</a> for your true peers. However, stay connected to DH, remain welcoming and open minded to the ideas and newbees that circulate in DH venues.</p>

<h2 id="is-the-quality-at-dh-venues-sinking">Is the quality at DH venues sinking?</h2>
<p>Many really good papers get rejected at DH conferences. Complaints about the quality of the papers that get in are rife. This is not all knee-jerk response from the authors that get rejected. Yes, getting a rejection on a good paper that you poured your heart and soul in, is an utter personal insult and just very hard to deal with in general. Trust me, there are forty years tenured full professors that get drunk the same night just to prevent themselves from buying a chainsaw to go primeval on the program committee. But to the outside world we maintain the civility that “rejection is part of academic life”. However the levels of good-but-rejected and bad-but-accepted do seem unsatisfying particularly in the case of DH conferences.</p>

<p>My personal perspective on this is that this is an inevitability for a growing community in an intrinsic interdisciplinary field. You cannot be expert in everything. I think I can just about manage vouching for a paper that attacks a literary problem in a not too advanced statistical manner. Let’s say I can judge the methodological and subject matter quality of a paper that deals with the rise of gender as a motive in Dutch literature from 1800CE to 2000CE. I think I would be able to evaluate the operationalization of that question. This alone involves many difficult problems. How does one establish that a motive of gender is actually present in a novel? And if it is present how is it counted? Do we count instances of gender motives into some categories? Do we take the number of words dedicated to the motive as a proxy of its prevalence? How to deal with the near sure possibility that 19th century authors were less open on this matter than 1980s authors? Do we acknowledge the fact that the subject has had many unheard voices or do we take what has been published as de facto “literature”? The problem is so drastically multifaceted that we will have years of test driving different operationalizations before we can even start thinking about truly taking our results as more than “interesting”. Oh… and this is the subject matter side. On the other hand I need to evaluate whatever the author decided to do with middle of the road statistics, or worse with fancy Bayesian cultural evolutionary models, or LLMs. I can actually check frequentist probabilities, I know how to interpret their numbers and am able to see if the calculations make sense. I can more or less follow what the Bayesian kool kids are doing, and I developed more or less a feeling for when their numbers do not really add up. I am as non-knowledgeable about LLMs as any of us currently.</p>

<p>And I am kind of an expert in my field. So people tell me, okay? I would never self-address myself as such. That is something dumb people do. That is all just to say that being a decent peer in a deeply interdisciplinary field is really really hard. As soon as an author is talking about, say, 12th century Chinese Buddhist manuscripts, I am totally lost on subject matter. All I can do is try to understand what the author wants to do regarding that subject matter, and assume that that research question is actually opportune and interesting for the people in the research community that is involved with the study of 12th century Chinese Buddhist manuscript. Operationalization: not a chance in hell my judgement is going to be expert level informed. Way too many confounding factors that I do not even know exist. Checking the statistics? Well, to the point that I will have to believe the counts of some form that arise from an operationalization I cannot judge that is applied to data of which I cannot evaluate if the curation is congruent with the research challenge. Bracket all that insecurity, and I can tell you if the researcher applied a cosine distance measure decently. If I can read the code, that is. Or, barring access to code, what the author says the techy did who was actually supervised by a statistician from the math department whom the author asked for advice and, chances are, who knows even less about manuscripts than I do.</p>

<p>You can kind of see, I hope, of how many factors and aspects one actually needs to have at least a bit of a grasp of to be a good peer reviewer in DH. Now, this problem scales. Put more domains in the meanwhile proverbial big tent and the chances of hitting a subject matter knowledgeable peer dwindle to parts of percentages.</p>

<h2 id="what-does-your-evidence-look-like-different-styles-of-science">What does your evidence look like? Different styles of science</h2>
<p>With domains and subdomains come particular conventions for method and style. The NLP people like the following. You start with a very concrete, very clear, very detailed problem description. Then you follow by a concise and clear data description. The next section describes the methods and techniques used. It is important (for some reason) that this section is undone from any, and I really mean <em>any</em>, fluff and detail that may be found in previous publications already, so that your less informed peer will have to look up at least twelve other publications to puzzle together what the bone dry skeleton idiomatic sentences interspersed with “[3,12,4,5]” and “[6,5]” mean. Now come the results, preferably as F1 measures listed together with how other approaches in the past performed. This is called evaluation. Your purpose in life as a peer reviewer coming from the NLP domain is to stomp down on any paper that does not have such a meticulous evaluation section. To kill it. To mock it. And to disqualify the whole community that enabled the existence of such a paper. Just saying, really. Lastly, you provide a bit of discussion and, if you really cannot help yourself, a future directions section. Done.</p>

<p>This is rather remote from someone arriving from philosophy whose paper as a first sentence has: “My project is to interrogate the tacit metaphysical presuppositions underwriting our ordinary grammar of agency, and thereby to delineate the conditions under which responsibility can be said to emerge as a normatively binding structure within shared forms of life.” My NLP peer will have ran screaming before even reaching the first comma. Yet the philosopher may  be making a completely valid contribution to a DH conference, because she is interested in how the subtleties of Wittgenstein’s thinking may speak to the methods we use for inquiring into language used in scientific discourse. This does not need stringent computational tractability and operationalization yet. It just needs decent examples that intrigue and challenge our current thinking. Because, you know, that is what we do in the humanities: create and test drives new intellectual ideas and reasoning, to see if they are palatable and lead anywhere interesting. Yep, that is very different from problem solving, and from task scoring boards hunting. It did however inspire centuries worth of new thought. If you are not interested in that as a computational linguist, that is fine. Just do not exhibit the hubris of some finer specimens I witnessed, saying that it is all just rubbish in DH because they did not recognize it as being research. That is your limited understanding of science speaking sir, not actual wisdom. Oops. Apologies if I sounded a bit… nasty there, that is because I have an opinion about that. Go read a book on the history of philosophy some time to value its actual pivotal importance.</p>

<p>No worries, everyone gets a wack with the same bat here. Plenty of stupidity and misunderstanding of how other fields operate in other specialists too. Because yes, the historian specializing in ancient Nubia who got into DH “because data” is very much at risk doing the same thing. He just read a marvelous argumentative contribution on the sense and nonsense of LLMs for historical research. It was enlightening, for in fact it was well written, it did point out some valuable do’s and don’ts, and it informed him how he could make use -responsibly and FAIR of course- of some LLM annotation for his data. A great timesaver found there. There, 90% relevance, solid contribution. But now… he is looking at this terribly densely written paper, with formulas and what have you. It seems to want to pack years of math learning into five hundred words, but it focuses on improving HTR with what? 1%? Why?! What use? HTR was last year! Okay, let’s write it is “solid work”, “a moderate contribution”, “more fit for a poster maybe then a long paper”. He is actually right, from his domain’s perspective, and at the same time he is “not even wrong”.</p>

<p>So it is all rather like too many Venn diagrams of domains with different research styles and different styles of communication. What counts as a valid research question and what counts as evidence in domain A will be vastly different from what counts as research and valid contribution in another. For all of the above, my reasoning is that any intrinsic interdisciplinary field will have more peer review misunderstanding than any sharply focused single subject or single method venue. So yes, the good-but-rejected paper ratio will be higher, so will be the bad-but-accepted paper ratio. Now, the hard questions are: is DH doing particularly worse than other interdisciplinary fields? And: is DH particularly harsh on more technical contributions?</p>

<h2 id="could-we-measure-it">Could we measure it?</h2>
<p>I do not have the answers to that. Yet. It would be interesting research wouldn’t it? Has anybody done this? Not so much for DH I think, but there might be other conferences that have their bibliographies perused for similar questions. The measure is not too hard to think up. It is just an F1, right?</p>

<div style="width:100%;margin-bottom:1em;">
<div style="width:50%;margin:auto;">
<svg style="vertical-align: -1.74ex;" xmlns="http://www.w3.org/2000/svg" width="27.994ex" height="4.889ex" viewBox="0 -1392 12373.1 2161" xmlns:xlink="http://www.w3.org/1999/xlink" aria-hidden="true">
  <defs>
    <style>
      svg a{fill:blue;stroke:blue}
      [data-mml-node="merror"]>g{fill:red;stroke:red}
      [data-mml-node="merror"]>rect[data-background]{fill:yellow;stroke:none}
      [data-frame],[data-line]{stroke-width:70px;fill:none}
      .mjx-dashed{stroke-dasharray:140}
      .mjx-dotted{stroke-linecap:round;stroke-dasharray:0,140}
      use[data-c]{stroke-width:3px}
    </style>
    <path id="MJX-254-NCM-I-1D439" d="M199 656C199 646 210 641 231 641C271 641 291 637 291 628C291 625 289 618 286 605L156 82C150 60 140 47 126 42C119 40 101 39 70 39C48 39 38 37 38 16C38 5 44 0 57 0L187 3L335 0C351 0 359 8 359 23C359 40 348 39 322 39C278 39 253 41 247 46C244 47 243 51 243 56L307 321L400 321C446 321 478 319 478 281C478 268 476 252 471 233C470 230 469 226 468 222C468 211 473 206 484 206C489 207 496 214 503 229L557 444C559 452 560 458 560 461C557 470 551 475 544 475C536 475 530 467 526 452C516 414 502 390 486 378C470 366 442 360 402 360L317 360L379 606C387 640 387 641 429 641L559 641C656 641 700 626 700 537C700 517 699 499 697 484C696 473 695 467 695 466C695 455 700 450 710 450C720 450 726 459 728 478L748 649C751 677 744 680 717 680L233 680C210 680 199 679 199 656Z" />
    <path id="MJX-254-NCM-N-31" d="M269 666C228 624 168 603 89 603L89 564C141 564 184 572 217 588L217 82C217 64 213 52 204 47C195 42 170 39 130 39L95 39L95 0C120 2 174 3 257 3C340 3 394 2 419 0L419 39L384 39C343 39 318 42 310 47C302 52 297 64 297 82L297 636C297 660 295 666 269 666Z" />
    <path id="MJX-254-NCM-I-1D445" d="M739 531C739 582 713 621 662 649C621 672 572 683 517 683L235 683C212 683 202 682 202 659C202 652 205 647 211 646C221 645 229 644 234 644C264 643 281 641 286 640C291 639 294 636 294 631C294 629 293 623 290 613L158 82C153 60 143 46 128 41C121 39 103 38 72 38C50 38 41 37 41 15C41 4 47-1 59 0L183 3L309 0C325-1 333 8 333 23C333 33 322 38 301 38C261 38 241 43 241 52C241 52 242 54 244 68L308 327L423 327C492 327 527 298 527 241C527 232 522 210 513 174C502 132 497 104 497 89C497 13 556-22 632-22C660-22 687-9 714 16C741 41 755 68 755 96C755 106 750 111 739 111C732 111 726 106 723 95C712 64 698 41 682 28C666 15 651 8 636 8C613 8 601 27 601 64C601 88 604 125 611 176C614 197 615 212 615 223C615 276 587 315 531 339C625 362 739 429 739 531M609 616C629 603 639 581 639 550C639 530 635 507 626 480C599 398 531 357 422 357L316 357L379 610C384 631 392 642 403 643C408 644 428 644 463 644C528 644 566 642 609 616Z" />
    <path id="MJX-254-NCM-N-2062" d="" />
    <path id="MJX-254-NCM-I-1D452" d="M124 129C124 153 129 186 139 227L188 227C253 227 303 235 339 250C372 264 394 284 405 309C412 326 415 342 415 355C415 410 363 442 307 442C268 442 229 432 190 412C113 372 46 281 46 171C46 69 105-11 204-11C257-11 304 2 345 27C379 48 404 69 420 90C427 99 430 106 430 109C430 120 425 126 414 126C409 126 404 122 398 114C365 70 324 42 277 30C246 22 223 18 206 18C149 18 124 72 124 129M375 355C375 289 311 256 182 256L147 256C166 322 194 366 232 387C262 404 287 413 307 413C343 413 375 391 375 355Z" />
    <path id="MJX-254-NCM-I-1D434" d="M137 3C155 3 217 0 236 0C251 0 258 8 258 23C258 32 252 38 239 39C210 40 196 50 196 69C196 78 205 98 224 128C251 173 270 206 283 228L526 228C526 221 527 207 530 185C537 112 541 72 541 67C541 48 519 39 474 39C455 39 446 31 446 15C446 5 452 0 464 0C488 0 564 3 588 3C608 3 679 0 699 0C714 0 721 8 721 24C721 34 712 39 694 39C666 39 649 41 644 44C639 47 635 56 634 71L574 689C571 708 572 716 551 716C539 716 529 710 522 697L178 120C148 70 108 43 59 39C43 38 35 30 35 15C35 5 41 0 52 0C68 0 121 3 137 3M492 577L522 267L307 267Z" />
    <path id="MJX-254-NCM-I-1D450" d="M328 325C328 300 341 287 368 287C404 287 427 318 427 354C427 410 368 442 308 442C239 442 176 412 122 353C68 294 41 229 41 159C41 61 106-11 204-11C257-11 304 2 345 27C379 48 404 69 420 90C427 99 430 106 430 109C430 120 425 126 414 126C409 126 404 122 398 114C365 70 324 42 277 30C246 22 223 18 206 18C149 18 121 53 121 122C121 184 150 274 174 317C198 361 249 413 308 413C346 413 372 402 386 381C355 378 328 358 328 325Z" />
    <path id="MJX-254-NCM-N-3D" d="M698 367L80 367C64 367 56 359 56 344C56 329 64 321 80 321L698 321C714 321 722 329 722 344C722 356 711 367 698 367M698 179L80 179C64 179 56 171 56 156C56 141 64 133 80 133L698 133C714 133 722 141 722 156C722 169 711 179 698 179Z" />
    <path id="MJX-254-NCM-N-32" d="M237 666C186 666 143 648 106 612C69 576 50 534 50 483C50 449 75 424 106 424C136 424 161 450 161 480C161 513 137 536 105 536C102 536 100 536 98 535C117 584 161 627 224 627C306 627 352 556 352 470C352 403 318 331 250 255L62 43C49 28 50 29 50 0L421 0L450 180L417 180C409 129 402 100 396 91C391 86 361 84 306 84L139 84L236 179C304 243 390 312 419 365C439 400 449 435 449 470C449 588 357 666 237 666Z" />
    <path id="MJX-254-NCM-N-22C5" d="M192 250C192 279 168 303 139 303C110 303 86 279 86 250C86 221 110 197 139 197C168 197 192 221 192 250Z" />
    <path id="MJX-254-NCM-I-1D43A" d="M324-22C412-22 481 5 532 59C538 46 564 1 578 1C583 1 586 4 588 8C590 12 597 32 606 68L624 144C630 167 634 184 637 195C648 239 648 238 702 239C715 239 721 247 721 263C721 273 716 278 705 278C686 278 620 274 601 275L462 278C446 278 438 270 438 254C438 245 444 241 456 240C513 237 543 235 546 233C549 231 550 227 550 222C550 215 543 185 530 133C510 61 436 17 343 17C221 17 148 98 148 220C148 239 150 263 153 290C164 372 218 487 261 540C311 603 403 666 505 666C609 666 660 587 660 479C660 470 657 438 657 429C657 420 663 415 676 415C681 415 685 416 688 417C693 424 696 431 698 438L760 691C760 700 755 705 745 705C741 705 735 701 727 692L662 619C623 676 568 705 497 705C442 705 388 692 333 667C222 615 141 533 89 421C63 366 50 310 50 253C50 93 164-22 324-22Z" />
    <path id="MJX-254-NCM-N-2061" d="" />
    <path id="MJX-254-NCM-N-2B" d="M698 274L413 274L413 559C413 575 405 583 389 583C373 583 365 575 365 559L365 274L80 274C64 274 56 266 56 250C56 234 64 226 80 226L365 226L365-59C365-75 373-83 389-83C405-83 413-75 413-59L413 226L698 226C714 226 722 234 722 250C722 263 711 274 698 274Z" />
    <path id="MJX-254-NCM-I-1D435" d="M756 543C756 634 666 683 568 683L236 683C213 683 203 682 203 659C203 643 223 644 235 644C265 643 283 641 288 640C293 639 295 636 295 631C295 629 294 623 291 613L159 82C154 60 144 47 129 42C122 40 104 39 73 39C52 39 42 31 42 15C42 5 52 0 73 0L426 0C491 0 551 20 608 59C671 102 703 155 703 217C703 296 638 344 567 358C655 379 756 446 756 543M554 644C623 644 658 612 658 547C658 496 638 454 596 420C554 386 507 370 456 370L317 370L377 610C385 644 385 644 427 644M582 299C596 276 603 253 603 228C603 176 583 131 542 94C501 57 455 39 402 39L267 39C257 39 250 39 246 40C240 40 237 42 237 45C237 46 237 49 238 53C269 180 293 276 309 340L493 340C536 340 565 326 582 299Z" />
  </defs>
  <g stroke="black" fill="black" stroke-width="0" transform="scale(0.7,-0.7)">
    <g data-mml-node="math">
      <g data-mml-node="msub">
        <g data-mml-node="mi">
          <use data-c="1D439" xlink:href="#MJX-254-NCM-I-1D439" />
        </g>
        <g data-mml-node="TeXAtom" transform="translate(676,-150) scale(0.707)" data-mjx-texclass="ORD">
          <g data-mml-node="msub">
            <g data-mml-node="mn">
              <use data-c="31" xlink:href="#MJX-254-NCM-N-31" />
            </g>
            <g data-mml-node="TeXAtom" transform="translate(533,-152.7) scale(0.707)" data-mjx-texclass="ORD">
              <g data-mml-node="mi">
                <use data-c="1D445" xlink:href="#MJX-254-NCM-I-1D445" />
              </g>
              <g data-mml-node="mo" transform="translate(759,0)">
                <use data-c="2062" xlink:href="#MJX-254-NCM-N-2062" />
              </g>
              <g data-mml-node="mi" transform="translate(759,0)">
                <use data-c="1D452" xlink:href="#MJX-254-NCM-I-1D452" />
              </g>
              <g data-mml-node="mo" transform="translate(1225,0)">
                <use data-c="2062" xlink:href="#MJX-254-NCM-N-2062" />
              </g>
              <g data-mml-node="mi" transform="translate(1225,0)">
                <use data-c="1D434" xlink:href="#MJX-254-NCM-I-1D434" />
              </g>
              <g data-mml-node="mo" transform="translate(1975,0)">
                <use data-c="2062" xlink:href="#MJX-254-NCM-N-2062" />
              </g>
              <g data-mml-node="mi" transform="translate(1975,0)">
                <use data-c="1D450" xlink:href="#MJX-254-NCM-I-1D450" />
              </g>
            </g>
          </g>
        </g>
      </g>
      <g data-mml-node="mo" transform="translate(2620,0)">
        <use data-c="3D" xlink:href="#MJX-254-NCM-N-3D" />
      </g>
      <g data-mml-node="TeXAtom" data-mjx-texclass="ORD" transform="translate(3675.8,0)">
        <g data-mml-node="mfrac">
          <g data-mml-node="mrow" transform="translate(2969.4,676)">
            <g data-mml-node="mn">
              <use data-c="32" xlink:href="#MJX-254-NCM-N-32" />
            </g>
            <g data-mml-node="mo" transform="translate(722.2,0)">
              <use data-c="22C5" xlink:href="#MJX-254-NCM-N-22C5" />
            </g>
            <g data-mml-node="mrow" transform="translate(1222.4,0)">
              <g data-mml-node="mi">
                <use data-c="1D43A" xlink:href="#MJX-254-NCM-I-1D43A" />
              </g>
              <g data-mml-node="mo" transform="translate(786,0)">
                <use data-c="2061" xlink:href="#MJX-254-NCM-N-2061" />
              </g>
              <g data-mml-node="mi" transform="translate(786,0)">
                <use data-c="1D434" xlink:href="#MJX-254-NCM-I-1D434" />
              </g>
            </g>
          </g>
          <g data-mml-node="mrow" transform="translate(220,-686)">
            <g data-mml-node="mrow">
              <g data-mml-node="mn">
                <use data-c="32" xlink:href="#MJX-254-NCM-N-32" />
              </g>
              <g data-mml-node="mo" transform="translate(722.2,0)">
                <use data-c="22C5" xlink:href="#MJX-254-NCM-N-22C5" />
              </g>
              <g data-mml-node="mrow" transform="translate(1222.4,0)">
                <g data-mml-node="mi">
                  <use data-c="1D43A" xlink:href="#MJX-254-NCM-I-1D43A" />
                </g>
                <g data-mml-node="mo" transform="translate(786,0)">
                  <use data-c="2061" xlink:href="#MJX-254-NCM-N-2061" />
                </g>
                <g data-mml-node="mi" transform="translate(786,0)">
                  <use data-c="1D434" xlink:href="#MJX-254-NCM-I-1D434" />
                </g>
              </g>
            </g>
            <g data-mml-node="mo" transform="translate(2980.7,0)">
              <use data-c="2B" xlink:href="#MJX-254-NCM-N-2B" />
            </g>
            <g data-mml-node="mrow" transform="translate(3980.9,0)">
              <g data-mml-node="mi">
                <use data-c="1D435" xlink:href="#MJX-254-NCM-I-1D435" />
              </g>
              <g data-mml-node="mo" transform="translate(759,0)">
                <use data-c="2062" xlink:href="#MJX-254-NCM-N-2062" />
              </g>
              <g data-mml-node="mi" transform="translate(759,0)">
                <use data-c="1D434" xlink:href="#MJX-254-NCM-I-1D434" />
              </g>
            </g>
            <g data-mml-node="mo" transform="translate(5712.1,0)">
              <use data-c="2B" xlink:href="#MJX-254-NCM-N-2B" />
            </g>
            <g data-mml-node="mrow" transform="translate(6712.3,0)">
              <g data-mml-node="mi">
                <use data-c="1D43A" xlink:href="#MJX-254-NCM-I-1D43A" />
              </g>
              <g data-mml-node="mo" transform="translate(786,0)">
                <use data-c="2061" xlink:href="#MJX-254-NCM-N-2061" />
              </g>
              <g data-mml-node="mi" transform="translate(786,0)">
                <use data-c="1D445" xlink:href="#MJX-254-NCM-I-1D445" />
              </g>
            </g>
          </g>
          <rect width="8457.3" height="60" x="120" y="220" />
        </g>
      </g>
    </g>
  </g>
</svg>
</div>
</div>

<pre style="font-size:40%;">
where:
GA = Good and Accepted
GR = Good but Rejected
BA = Bad but Accepted
</pre>

<p>That will give you a harmonic mean telling you how good DH is at selecting ‘good’ papers. The harder part is figuring out what the “good-but-rejected” and “bad-but-accepted” papers are. We cannot just take the word of the authors for it. Or well we actually could. We could just do that survey. We could do an appeal to honesty and ask if they themselves thought it really should have been accepted, also because they can point to a later paper that was accepted elsewhere. See what those numbers tell us. But it will require some thought on how to research this well from a bibliometrics point of view. There must be ways however, I am sure. We can probably do some Bayesian statistics, maybe even just based on keywords and titles of papers, see what becomes more probable for rejection or acceptance over time, controlling for unavoidable buzzword fads. Something such. Maybe also a simulation? Any takers yet?</p>

<h2 id="open-minds-generous-research">Open minds, generous research</h2>
<p>The particular unease that we feel about having been rejected, or accepted for that matter, to DH, has driven many of us towards CHR and CCLS. In a sense this is a good thing. It is good that there is broad DH and deep CH. At least my personal take on that is: it is good that there is some sort of entry level where curious minds can come into the field of DH and CH. This is what ADHO DH and <a href="https://eadh.org/news/2025/12/04/cfp-eadh-conference-2026-linking-europe-digital-humanities-without-borders">EADH</a> are good for. Plus, DH is the community motor around the world. I think hardliners in CHR and CCLS should be wiser than to just ridicule the work that is put into these venues. Yes, sometimes not too clever things are done there, but people coming from a humanities background cannot go from zero to 100% full on computational savvy in a matter of a year. Big tent venues like DH have a role here. CHR and CCLS should actually make sure they keep their ties to DH venues warm. We all need these connections to be strong to make sure that the right experts find each other. Being an expert is easy. Connecting expertise, fostering a community for the better of science is hard, and mostly still not seen as a core contribution.</p>

<p>I think the worst thing we can do is feed some animosity between venues and subdomains. But it is hard to not walk into that pitfall. “Another rejection! And see here, that awful paper got accepted. F*** it. DH just sucks. I will only submit to x, y, z in the future.” Backing into our expertise and niche corners, sulking, is narrow minded. However, we really need broad-minded, and generosity. Of subject, of method, of heart. Be an expert. Then be generous with your expertise.</p>

<p>–JZ_20260304_2012</p>

<h2 id="postscript">Postscript</h2>
<p>This post grew out of a tongue in cheek thinking up of peer review scenarios gone wonky. By genuine good intent and being all expert professional we create our own rejection hell. I think these things genuinely happen. They drive the dynamics I wrote about in the post.</p>

<p><em>Peer incompatibility</em>
Suppose the paper is highly technical, but the peer is a theorist; or vice versa. The theorist reviewer will point out practical problems that are non-pertinent to the specifics of the use case the technologists present, like that it does not solve one very obscure very different problem that he witnessed in 1996. Peer will ask for more theoretical framing. Technical peer will complain that there is no application, let alone any evaluation, and that the paper comes across as muddled and confused anyway.</p>

<p><em>Peer Dunning-Kruger effect</em>
People susceptible to this effect will generally tick the box “My expertise, I am highly knowledgeable about this subject”, while true experts usually are more modest. Reviews will contain things like “the authors have ignored [this or that] body of research” or “there is a whole field dedicated to this problem, why didn’t the authors…” or, the worst: “this is a solved problem”. The level of brutality in the review is usually reversely proportional to the lack of knowledge the expert exhibits. (This last sentence requires thinking.)</p>

<p><em>Domain reporting style mismatch</em>
This is where NLP/STEM ‘problem-data-method-results-discussion’ style papers or expectations meet humanities ‘why-this-proposed-solution-actually-is-a-problem’ style papers or expectations. A humanist peer will likely say “Interesting subject, but I think this would be more fit as a short paper, or rather better even a poster”. The technical peer will point out that there is not even a research question in the paper, and that it in general seems a non-problem that is being discussed which the author could have known if he had read x, y, and z in the NLP domain literature.</p>

<p><em>Subject matter SOTA (state of the art) mismatch</em>
When the technical solution is highly appreciated, but non of the technical peers actually succeeded in pointing out that we literally know what happened at Waterloo basically second by second from historical sources and re-enactment and that therefore the 14 million Euros spent on the simulation with 100,000 individually modeled computational agents, which also required an expensive industry game studio to recreate the battle field and physics, actually does not add all that much.</p>

<p><em>Technology SOTA mismatch</em>
When the historians go absolutely wild and wanking because someone made a nice GUI and a model that you can ask “How long did it take to get from Colonia (Cologne) to Roma on horseback?” Nobody thinks it weird that the answer it comes up with is an averaged 800 hours or a 6 months turnaround, while we know that it would take a mere three days if the message was really important. The one peer that had statistical and topological knowledge tried to point out that the network model used basically added up to a stochastic modeling, but none of the PC understood what she was pointing out.</p>

<p>–</p>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[TL;DR You have been rejected at DH. This may well be because you are more expert than many DH peers and your work is more advanced in a very particular niche. There is no hard data on this type of rejection, but it would be interesting to figure out how to do the bibliometrics and actually do them. Anyway, do not sulk, do not mock. Submit to satellite conferences such as CHR and CCLS for your true peers. However, stay connected to DH, remain welcoming and open minded to the ideas and newbees that circulate in DH venues.]]></summary></entry><entry><title type="html">Who Cares Whence the Words?</title><link href="http://localhost:4000/who-cares-whence-the-words/" rel="alternate" type="text/html" title="Who Cares Whence the Words?" /><published>2023-04-20T15:52:56+02:00</published><updated>2023-04-20T15:52:56+02:00</updated><id>http://localhost:4000/who-cares-whence-the-words</id><content type="html" xml:base="http://localhost:4000/who-cares-whence-the-words/"><![CDATA[<p>A lot of hubbub and fuzz over a <a href="https://www.nature.com/articles/d41586-023-00107-z">few authors listing ChatGPT as author on articles in scientific journals</a> (<a name="ref_walker_2023"><a href="#walker_2023">Stokel-Walker 2012</a>). On closer inspection nothing really dramatic is going on.</a></p>

<p>The <em>Nature</em> news item lists four known cases, two of which were editorial process oversights. We can ignore those. There was one that looked into GPT’s capacity to write about itself. This is, I would argue, category error. Listing excerpts from what GPT writes about itself is quoting. Normally you would refer to the publication you quote from, and in this case one would need to refer to the date of use, version, prompt, and settings. Properly citing a GPTChat prompt is probably a rabbit hole by itself, but one could and should. Listing GPT as an author in such a case is like listing Umberto Eco as an author when you quote his <a name="ref_eco_1981"><a href="#eco_1981"><em>Theory of Signs and the Role of the Reader</em></a></a> (1981) at length in your article. That is just not right and not proper. One should cite and not appropriate authorship in that case. No doubt the whole category play is intentional and fun, but in the end it is also a bit.. well foolish and whimsical.</p>

<p>The one that really would make me go jeeps, is the one where researchers use ChatGPT to generate a two pager on the pros and cons of using <a href="https://doi.org/10.18632%2Foncoscience.571">Rapamycin in the context of the philosophical concept of Pascal’s wager</a> (<a name="ref_zhavoronkov_2022"><a href="#zhavoronkov_2022">Zhavoronkov 2022</a></a>). In this case one still gets the impression that the primary investigators have checked what ChatGPT was generating, and there still seems a distinction between what ChatGPT and what the researchers contributed. But it also comes eerily close to “just letting ChatGPT have a go and see if the peer reviewers catch it”. That is not using ChatGPT as a writing aid but abusing it as a content generator which would become irresponsible and unethical as soon as not every word has been checked by the submitting researchers.</p>

<p>The key here is scientific responsibility of course. As long as the submitting authors have thoroughly checked, have agreed, and have confirmed that what has been written is truthfully representing what they researched, how they researched it, and what they found… who cares who generated the words? Much of the genuine writing by hard sciences researcher is so poor quality as to readability that, by all means, have ChatGPT have a go at it. But as soon as you generate a full paper, claim you have done the research but didn’t, it is fraud clearly.</p>

<p>Is my thinking anyway. ChatGPT<sup><a name="ref_note_001"><a href="#note_001">1</a></a></sup> had something so say too, obviously…</p>

<blockquote>
  <p>As ChatGPT, I believe that the topic of ChatGPT being listed as an author on scientific articles raises questions around agency and scientific responsibility. While ChatGPT is capable of generating responses to certain prompts, it lacks true agency in decision-making and cannot have ownership of its contributions. Therefore, it is important for humans to take full responsibility for the research and be transparent about the role of ChatGPT in the publication process.<br />
<br />
At the same time, scientists have a responsibility to consider the ethical implications of using AI-generated data to draw conclusions and make decisions. It is crucial that researchers ensure the validity and accuracy of the data generated by AI systems before using it for scientific purposes. This includes examining potential biases in the dataset, as well as addressing related issues such as privacy concerns and data ownership. Ultimately, maintaining a high standard of scientific responsibility will be vital in ensuring that AI-generated research is used safely and effectively.</p>
</blockquote>

<p>–JZ_20230420_1704</p>

<h3 id="notes">Notes</h3>

<p><span style="font-size:0.75em">[<a name="note_001"><a href="#ref_note_001">1</a></a>] This text from ChatGPT was generated through the <em>ChatGPT - Genie AI</em> plugin for <em>VSCode</em> on Thursday 20 April, 16.36CEST, using the gpt-3.5-turbo model with the prompt “Write an opinion on how everybody is having opinions on how ChatGPT is being listed as an author on scientific articles,” followed by the prompt “Do the same but include a sentence on agency and one on scientific responsibility.”</span></p>

<h3 id="references">References</h3>

<ul>
  <li><a name="walker_2023"><a href="#ref_walker_2023">Stokel-Walker, C. (2023).</a> <em>ChatGPT listed as author on research papers: many scientists disapprove</em>. Nature – News, 613: 620–21. <a href="https://doi.org/10.1038/d41586-023-00107-z">https://doi.org/10.1038/d41586-023-00107-z</a>.</a></li>
  <li><a name="eco_1981"><a href="#ref_eco_1981">Eco, U. (1981).</a> <em>The Theory of Signs and the Role of the Reader</em>. The Bulletin of the Midwest Modern Language Association, 14(1): 35–45. <a href="https://doi.org/10.2307/1314865">10.2307/1314865</a>.</a></li>
  <li><a name="zhavoronkov_2022"><a href="#ref_zhavoronkov_2022">Zhavoronkov, A</a></a>. (2022). <em>Rapamycin in the context of Pascal’s Wager: generative pre-trained transformer perspective</em>. Oncoscience, 9: 82–84. <a href="https://doi.org/10.18632/oncoscience.571">10.18632/oncoscience.571</a>.</li>
</ul>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[A lot of hubbub and fuzz over a few authors listing ChatGPT as author on articles in scientific journals (Stokel-Walker 2012). On closer inspection nothing really dramatic is going on.]]></summary></entry><entry><title type="html">A Tiny Cartography of Mapping Literature</title><link href="http://localhost:4000/a-tiny-cartography-of-mapping-literature/" rel="alternate" type="text/html" title="A Tiny Cartography of Mapping Literature" /><published>2022-12-01T13:58:00+01:00</published><updated>2022-12-04T14:29:00+01:00</updated><id>http://localhost:4000/a-tiny-cartography-of-mapping-literature</id><content type="html" xml:base="http://localhost:4000/a-tiny-cartography-of-mapping-literature/"><![CDATA[<p>My guess is that almost any reader will recognize the map from figure 1, just below. It is a map of Middle Earth. This image was reproduced from the 1993 paperback edition of J.R.R. Tolkien’s <em>The Lord of the Rings</em>. The map and its various siblings that have adorned some of the first pages of virtually every edition are pretty much emblematic for the archetypal use of maps in fiction, which seems to be to provide some cartography for the story world of the novel. Maps are typically embedded as front or back matter, but sometimes they are even independent companions to the primary work (fig. 2). Maps as paratext may relate to the real world, to imaginary geographies, and really anything in between.</p>

<p><img src="/assets/uploads/2022/middle_earth.jpg" alt="" />
<em>Fig. 1: Map of a fictional world.</em></p>

<p><img src="/assets/uploads/2022/discworld_map.jpg" alt="" />
<em>Fig. 2: The map of the Disc World.</em></p>

<p>Maps have played such a prominent role in literature that they now even need to be mapped out themselves, leading to interesting <a href="https://www.theguardian.com/books/2022/jan/19/mapping-fictions-relationship-authors-literary-maps">exhibitions</a> for instance. Apparently readers (or at least some readers) are fascinated by the geography of novel space. But their fascination is not limited to maps pictured in novels. Readers want to know <a href="https://www.groene.nl/artikel/kaart-van-nederlandse-schrijvershuizen">where authors lived</a> (fig. 3). They love “book lovers’ tours”, such as the one in <a href="https://www.bookloverstours.nl/">Amsterdam</a> (fig. 4), where many stories were situated. But there are plenty of walks <a href="https://maastrichtboekenstad.nl/literaire-plekjes/">outside of Amsterdam</a> as well. In fact, so many that they warrant a very <a href="https://indevoetsporenvanschrijvers.nl/">own website</a> (fig. 5) and <a href="https://play.google.com/store/apps/details?id=seven.client.letterkundig.android&amp;hl=nl&amp;pli=1">app</a> – although sadly, that last one seems out of service currently.</p>

<p><img src="/assets/uploads/2022/schrijvershuizen.png" alt="" />
<em>Fig. 3: Map of homes of Dutch literary authors.</em>
<img src="/assets/uploads/2022/bookloverstour.png" alt="" />
<em>Fig. 4: Amsterdam’s Book Lovers’ Tour.</em>
<img src="/assets/uploads/2022/voetsporen.png" alt="" />
<em>Fig. 5: Many more literary walks exist.</em></p>

<p>Others have creatively tried to map literary quotes or titles of novels. There is a map of Amsterdam made up of literary quotes for instance (fig. 6). And there is a map of the whole of <a href="https://www.halcyonmaps.com/#/map-of-the-literature/">world literature</a> (fig. 7) – obviously biassed and painfully selective as such a project must be.</p>

<p><img src="/assets/uploads/2022/amsterdam-in-quotes.jpg" alt="" />
<em>Fig. 6: Amsterdam in literary quotes.</em>
<img src="/assets/uploads/2022/map_world_literature.png" alt="" />
<em>Fig. 7: Martin Vargic, a map of world literature.</em></p>

<p>Notwithstanding that there clearly is a relation of interest between literary texts and space –be it narrative space or geographical space– the cartography of Dutch literature as a scientific practice seems to be at most in a nascent state. This may be related to the question what exactly should be or can be mapped. Certainly in literature all things are in principle fictional. Although any story may point to some reality that exists externally to it, we can never be sure about what we get told about that reality. How truthful or deceitful are the connections and relations between narrative and reality that the story wants us to believe that exist? It is of course exactly their fictionality that allows texts themselves to be maps that help us navigate the world we live in. Texts and narratives help us to understand and grapple with people, events, history, and society. They are traveling guides. But, as we know well, “the map is not the territory” (<a href="#korz-1958">Korzybsk 1958[1933]</a>). And therefore: how useful is it to paint a map of a story world? It might just suggest too much relevance of real world geography and physics for the narrative space, for at least in stories time travel is possible and roads need not to lead from A to B. The essence of the story space is probably exactly in how it deviates from the reality we perceive.</p>

<p>But narrative philosophy aside, there are also more practical matters. How does one map narrative space, or for that matter, literature? Cartography is a much more practiced exercise in linguistics. I am reminded of Margit Rem’s painstaking cartography of the language differences in the written diplomatic legacy of the court of Holland (fig. 8). Rem worked at the <a href="https://meertens.knaw.nl/">Meertens Institute</a>, which itself derives part of its fame from its long standing history of making maps of language, language change, and dialects, and so forth. Traditionally these were visual maps only, but by courtesy of new media also <a href="https://www.meertens.knaw.nl/projecten/sprekende_kaart/svg/">speaking maps</a> now exist.</p>

<p><img src="/assets/uploads/2022/klerken.png" alt="" />
<em>Fig. 8: An example of mapping the language of clerks working at the court of Holland in late medieval times. Reproduced from <a href="#rem_2003">Rem 2003</a>.</em></p>

<p>However, because language is strongly geographically linked, the role of cartography in linguistics is more or less obvious. In literary research that role is far less obvious and the relation between text and fictional geography is –often intentionally– complicated. Consequently, I think, we find few attempts so far to “map” Dutch literature. Herman Lodewick –famous for two text book anthologies that introduced several generations of secondary school pupils and students to the landscape of Dutch literature– published an “atlas” of Dutch literature together with two co-authors (fig. 9). In reality it was not so much a set of maps as it was “just another” anthology. Mapping Dutch literature has mostly meant describing a literary history, with sometimes a little attention to the relation between that literature and the <a href="https://www.literatuurgeschiedenis.org/middeleeuwen/de-aarde-is-een-bol">geographical world or world view it functioned within</a>. The Digital Library of Dutch Literature (<a href="https://www.dbnl.org/">DBNL</a>) seems to have gotten the farthest with an actual attempt at mapping Dutch literature, providing <a href="https://www.dbnl.org/atlas/nevl_standaard.php">a clickable map that ties author names to geographical locations</a>.</p>

<p><img src="/assets/uploads/2022/ik_probeer_mijn_pen.jpg" alt="" />
<em>Fig. 9: Lodewick’s “Atlas” of Dutch Literature.</em></p>

<p>With more and more literature becoming available as digital data, and with developing computational techniques and skill, this nascent state may be changing a little. One example of trying to computationally map a narrative world we find in <a href="#louwerse_2012">Louwerse and Benesh 2012</a>. The authors are interested in how spatial mental representations can be gauged from language, veering close towards the computation of fictional world maps, resulting in fascinating calculated maps (fig. 10).</p>

<p><img src="/assets/uploads/2022/louwerse_benesh_fig.gif" alt="" /></p>

<p>Mapping literature at large remains a desideratum. In the <em>Impact and Fiction</em> project we are most certainly not aiming to change or even add to the ground-works of digital literature cartography. However, while investigating the texts and metadata of our corpus of 19,622 works of fiction and non-fiction, we may serendipitously be producing some possibly interesting maps of Dutch literature, or part thereof. We are trying to make sense of what kind of works of fiction cause which kind of impact in what type of reader. To do this we need to be able to relate the vocabulary of the novels to impact described by readers in online reviews. In practice this means trying to reduce the countless linguistic and semantic features of the texts to a few that we hypothesize may relate to such reader impact. One notion of such features are the words that relate to topic, theme, or genre. Neither of those three concepts, however, has been unambiguously formalized and computationally operationalized. Nobody seems to know exactly what theme is, what topic looks like, and what a genre’s vocabulary is. These things are vague, and shifty, and hard to compute.</p>

<p>Nevertheless we try, and one attempt we did was in topic modeling the whole corpus. <a href="https://cbail.github.io/SICSS_Topic_Modeling.html">Topic modeling</a> tries to identify which words occur often together in texts, and subsequently clusters texts based on those shared occurrences. Effectively it “maps” texts to certain points in a very high dimensional space, and texts that use common word pairs often all end up somewhere close to the same point in that space. Unfortunately the 600,000 or so dimensional world is not easy for us to navigate as lower dimensional beings. We need techniques –in our case called <a href="https://umap-learn.readthedocs.io/en/latest/">UMAP</a>– to project that territory on to a two dimensional surface, just as cartographers do with three dimensional territories.</p>

<p>In the case of our corpus we end up with something that we could call a map of 19,000+ (fig. 11) novels. The colored clusters you see are clusters of topics and we were able to establish that these topics correlated very narrowly with genre. But another observation is that the clusters are also very much related to authors; see figure 12, where colors represent (a few) authors rather than topics. Does that mean that topics are both indicative for author and genre? Well, yes and no. Yes, because intuitively we can understand that many topics relate to specific genres (e.g. physical violence will be present more in war novels and murder mysteries, while less so maybe in a literary novellas). Also intuitively we can see how many authors will stay well within one (or a few) genres. But no, the other possibility that we need to delve into is that the selection of features we use and the techniques we apply are not sensitive enough to paint a decent map of the territory.</p>

<p><img src="/assets/uploads/2022/topic_clustering_nur_colors_legend.png" alt="" />
<em>Fig. 11: UMAP 2D projection of a high dimensional topic space, clearly showing how genres (colors) cluster.</em>
<img src="/assets/uploads/2022/topic_clustering_author_colors_legend.png" alt="" />
<em>Fig. 12: The same, but now colors have been used to indicated the most prolific authors.</em></p>

<p>Mapping Dutch literature using topic models results in a map that tells us how topics are distributed over genres and authors. But we learned from this that our cartographic tools are crude to say the least. We need better digital theodolites and a more sophisticated computational literary GPS. That is part of the work ongoing in our project. Of course there is no reason why we would not meanwhile enjoy the maps we already made.</p>

<h3 id="note">Note</h3>
<p>This blogpost is a copy of the one I wrote for the blog of <a href="https://impactandfiction.huygens.knaw.nl/">Impact &amp; Fiction</a> in december 2022.</p>

<h3 id="references">References</h3>

<ul>
  <li><a name="louwerse_2012">Louwerse, Max M., and Nick Benesh. 2012.</a> “Representing Spatial Structure Through Maps and Language: Lord of the Rings Encodes the Spatial Structure of Middle Earth.” <em>Cognitive Science</em> 36 (8): 1556–69. <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/cogs.12000">https://doi.org/10.1111/cogs.12000</a></li>
  <li><a name="korz-1958">Korzybski, Alfred. 1958.</a> <em>Science and Sanity: An Introduction to Non-Aristotelian Systems and General Semantics.</em> 5th (first published 1933). New York: International Non-Aristotelian / Institute of General Semantics.</li>
  <li><a name="rem_2003">M. Rem, De Taal van de Klerken Uit de Hollandse Grafelijke Kanselarij (1300-1340). 2003.</a> <em>Naar Een Lokaliseringsprocedure Voor Het Veertiende-Eeuws Middelnederlands</em>. Amsterdam: Stichting Neerlandistiek VU.</li>
</ul>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[My guess is that almost any reader will recognize the map from figure 1, just below. It is a map of Middle Earth. This image was reproduced from the 1993 paperback edition of J.R.R. Tolkien’s The Lord of the Rings. The map and its various siblings that have adorned some of the first pages of virtually every edition are pretty much emblematic for the archetypal use of maps in fiction, which seems to be to provide some cartography for the story world of the novel. Maps are typically embedded as front or back matter, but sometimes they are even independent companions to the primary work (fig. 2). Maps as paratext may relate to the real world, to imaginary geographies, and really anything in between.]]></summary></entry><entry><title type="html">On Code Literacy</title><link href="http://localhost:4000/on-code-literacy/" rel="alternate" type="text/html" title="On Code Literacy" /><published>2019-09-15T21:07:05+02:00</published><updated>2020-09-25T09:43:04+02:00</updated><id>http://localhost:4000/on-code-literacy</id><content type="html" xml:base="http://localhost:4000/on-code-literacy/"><![CDATA[<p>The following is the transcript of a short speech I contributed to the panel “Programming Humanists - What is the role of coding literacy in DH and why does it matter?” at the DH Benelux 2019 Conference, held at ULiège in Luik, Belgium. The organizers asked me to try to be provocative. That is, provocative to non coding humanists. However, the title of the panel – as I was afraid – caused a rather self-selecting crowd of code savvy scholars and engineers to turn up. So the statement itself was about as provocative as a tea to an Englishman. The panel in all was highly enjoyable and the panelists iterated their thoughts on how code literacy could be raised among humanists. I’ll leave the post here, and who knows, maybe some scholars will still be provoked by it…</p>

<h2 id="mea-culpa">Mea Culpa</h2>
<p>To start I quote Stephen Ramsay from when he was faced with a similar challenge: “[The organisers have] asked that we spend exactly [five] minutes giving our thoughts on this subject, and I like that a lot. With only [five] minutes, there’s no way you can get your point across while at the same time defining your terms, allowing for alternative viewpoints, or making obsequious noises about the prior work of your esteemed colleagues. Really, you can’t do much of anything except piss off half the people in the room.” End of quote. This is from his 2011 MLA contribution to the panel “History and Future of Digital Humanities”. He said the following quite controversial – as it turns out – thing: “Do you have to know how to code? I’m a tenured professor of digital humanities and I say «yes».” End of quote. So if I have pissed you off with this already, I have done my job. Otherwise sit tight and see if I can do better.</p>

<h2 id="why-it-matters">Why it matters</h2>
<p>Code literacy matters because in the near future, but even now already, veritably all information, cultural artefacts, and scholarly objects of study will be digital or will have digital aspects. This is tied to what science fiction writer Bill Gibson (2009) called the Eversion of Cyber Space, which happens if digital environment, human society, and even physical world become inextricably intertwined. This already happened according to Steven E. Jones. I agree, I even know the exact data: 29 June 2007. Anyone? It was the date of the introduction of the iPhone. The ubiquitous use of smartphones, Internet, the Web, digital tracking and surveillance techniques, digital streams, etc. etc. has made what once was a virtual world and different, a very normal part of reality indeed.</p>

<p>As a scholar in the humanities this impacts you, you cannot work with this digital information and these digital objects if you do not know how they work and how they get produced, which  always  involves code. To give but one example: how will you ever adequately describe, edit, and/or analyze J.R. Carpenter’s <a href="http://luckysoap.com/cityfish/"><em>CityFish</em></a> if you cannot read code? <em>CityFish</em> is a webbased literary creation that produces different text through JavaScript code.</p>

<p>It also matters because code is a tool of power to control and discipline. Annette Vee (2013) draws a parallel with the increased control government exercised in history by using administration and law rules, which require script, and thus writing and reading. Governments and industries are using code increasingly to enforce their power and further their interest. It is one of the roles of critical academics to study and interrogate these processes of power in society. But you cannot do this if you do not know how code works.</p>

<h2 id="what-do-we-mean-by-coding-literacy">What do we mean by “coding literacy”?</h2>
<p>And this is what we – or at least what I – mean by code literacy: the ability to know how code works. I follow Vee’s thinking and defining in this. Vee: “But, unfortunately, when «literacy» is connected to programming, it is often in unsophisticated ways: literacy as limited to reading and writing text; literacy divorced from social or historical context; literacy as an unmitigated form of progress.” Vee argues that almost every skill nowadays is called literacy in service of “urgency” (e.g. “quantitative literacy”). But we should be more careful when define what we mean… Vee defines “literacy” “as a human facility with a symbolic and infrastructural technology—such as a textual writing system—that can be used for creative, communicative and rhetorical purposes.” This is different than “material intelligence”, which is technology dependent communicative skill, or for short: being able to send an email. Material intelligence is a <em>nice to have</em>, literacy is a <em>must have</em>. Literacy is essential and required to navigate your world, which implies that the connected technology must be pervasive and central (or <em>infrastructural</em>) to a society. And yes, we just established that code is on its way there.</p>

<p>Literacy, according to me, and according to Alan Kay (1984), the inventor of SmallTalk, also implies fluency. That is the ability to work comfortably with a coding technology, and to be able to think, argue, and express yourself comfortably with larger, more meaningful chunks of expressions in such a literacy. But we all have to start somewhere, so copy paste coding is really fine, and it is okay for scholarly software to suck (Baldridge 2015).</p>

<h2 id="do-you-need-to-know-how-to-code">Do you need to know how to code?</h2>
<p>So if you ask me, as the organizers of this panel did, if “understanding code [is] an essential part of thorough DH scholarship?” The answer is: yes!</p>

<p><img src="/assets/uploads/2019/car-1024x530.jpg" alt="" />
<em>Fig. 1: A conventional metaphor used to deny the need for code literacy</em></p>

<p>We need to get rid of the pernicious metaphor of the car. That is: you will often hear people say something like “I don’t need to know how a car works to be able to drive in it; so I do not need to know code to understand a computer.” At the scholarly level this metaphor is wrong and it should be forbidden to be used.</p>

<p>It is nonsense that you do not need to know how a car works to be able to drive it. You do have tacit knowledge of how a car works, what is the front end, what is the back end. You do know how to operate the rather intricate set of levers, buttons, and wheels that make it go, and you know which lever causes what light to blink. And you know where the petrol goes, and that it doesn’t like gasoline. There is actually a lot more technical knowledge to driving than the car-computer metaphor suggests. And you get actual training for that.</p>

<p><img src="/assets/uploads/2019/crash-1024x530.jpg" alt="" />
<em>Fig. 2: But it is a faulty and pernicious metaphor</em></p>

<p>You would crash and burn without that training. If you are going to use code for your research you are going to crash and burn if you cannot understand at an adequate level what that code does. Partly because it is tied to methods and techniques of which you need to know what impact they have on your data, your analysis. And partly because the code and the techniques are made by other persons who built in assumptions that may be both benign and correct, but may also be malicious and faulty, given your context.</p>

<p>The better metaphor for trying to apply code objects without knowing code is trying to analyze German literature purely on the basis of English translations. You can, but you won’t be adequate, only vaguely right, and mostly producing cargo cult analyses.</p>

<h2 id="so-yes-you-need-code-literacy">So yes, you need code literacy</h2>

<p>Thus I am happy to now play on a provocative statement made by Peter Robinson during the 2013 ADHO Digital Humanities Conference in Lincoln, Nebraska. He said: “Digital humanists should get out of textual scholarship: and if they will not, textual scholars should throw them out.” I never liked that statement much. I will rephrase it more appropriately now: “Scholars that cannot code should get with the program, and if they will not, code literate scholars should throw them out.”</p>

<p>–JZ_20190910_1447</p>

<h3 id="references">References</h3>

<ul>
  <li>Baldridge, J. (2015) ‘It’s okay for academic software to suck’, Java Code Geeks, 12 May. Available at: <a href="https://www.javacodegeeks.com/2015/05/its-okay-for-academic-software-to-suck.html">https://www.javacodegeeks.com/2015/05/its-okay-for-academic-software-to-suck.html</a> (Accessed: 25 April 2016).</li>
  <li>Carpenter, J. R. (2010) <em>CityFish</em>, J. R. Carpenter || Luckysoap &amp; Co. Available at: <a href="http://luckysoap.com/cityfish/">http://luckysoap.com/cityfish/</a> (Accessed: 7 June 2017).</li>
  <li>Estrada, M., Liliana, Wigham, M. and Koolen, M. (2019) ‘Programming humanists - What is the role of coding literacy in DH and why does it matter?’, in <em>DH Benelux 2019</em>. DH Benelux 2019, Liège: Université de Liège, p. 25. Available at: <a href="http://2019.dhbenelux.org/wp-content/uploads/sites/13/2019/08/DH_Benelux_2019_paper_25.pdf">http://2019.dhbenelux.org/wp-content/uploads/sites/13/2019/08/DH_Benelux_2019_paper_25.pdf</a> (Accessed: 15 September 2019).</li>
  <li>Gibson, W. (2009) <em>Spook Country</em>. Reprint edition. New York: The Berkley Publishing Group (Blue Ant, Book 2).</li>
  <li>Jones, S. E. (2014) <em>The Emergence of the Digital Humanities</em>. New York, London: Routledge.</li>
  <li>Kay, A. (1984) ‘Computer Software’, <em>Scientific American</em>, September, pp. 53–59. Available at: <a href="https://www.nature.com/scientificamerican/journal/v251/n3/pdf/scientificamerican0984-52.pdf">https://www.nature.com/scientificamerican/journal/v251/n3/pdf/scientificamerican0984-52.pdf</a> (Accessed: 14 March 2018).</li>
  <li>Ramsay, S. (2011) ‘Who’s In and Who’s Out’, Stephen Ramsay — Blog, 8 January. Available at: <a href="https://web.archive.org/web/20170721063833/http://stephenramsay.us/text/2011/01/08/whos-in-and-whos-out/">https://web.archive.org/web/20170721063833/http://stephenramsay.us/text/2011/01/08/whos-in-and-whos-out/</a> (Accessed: 21 July 2017).</li>
  <li>Robinson, P. (2013) ‘Why digital humanists should get out of textual scholarship. And if they don’t, why we textual scholars should throw them out.’, Scholarly Digital Editions, 29 July. Available at: <a href="http://scholarlydigitaleditions.blogspot.nl/2013/07/why-digital-humanists-should-get-out-of.html">http://scholarlydigitaleditions.blogspot.nl/2013/07/why-digital-humanists-should-get-out-of.html</a> (Accessed: 10 January 2018).</li>
  <li>Vee, A. (2013) ‘Understanding Computer Programming as a Literacy’, <em>LiCS,</em> 1(2), pp. 42–64. Available at: <a href="http://licsjournal.org/OJS/index.php/LiCS/article/view/24/26">http://licsjournal.org/OJS/index.php/LiCS/article/view/24/26</a> (Accessed: 24 February 2014).</li>
</ul>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[The following is the transcript of a short speech I contributed to the panel “Programming Humanists - What is the role of coding literacy in DH and why does it matter?” at the DH Benelux 2019 Conference, held at ULiège in Luik, Belgium. The organizers asked me to try to be provocative. That is, provocative to non coding humanists. However, the title of the panel – as I was afraid – caused a rather self-selecting crowd of code savvy scholars and engineers to turn up. So the statement itself was about as provocative as a tea to an Englishman. The panel in all was highly enjoyable and the panelists iterated their thoughts on how code literacy could be raised among humanists. I’ll leave the post here, and who knows, maybe some scholars will still be provoked by it…]]></summary></entry><entry><title type="html">A Note on Interpretation</title><link href="http://localhost:4000/a-note-on-interpretation/" rel="alternate" type="text/html" title="A Note on Interpretation" /><published>2019-07-22T13:21:06+02:00</published><updated>2020-09-25T12:14:14+02:00</updated><id>http://localhost:4000/a-note-on-interpretation</id><content type="html" xml:base="http://localhost:4000/a-note-on-interpretation/"><![CDATA[<p>Thinking and reasoning about what interpretation exactly is, is an endless source of joyful wondering. I was just rereading my own <em>Screwmeneutics</em> (2016) – while preparing my thesis conclusion, there was no vanity in that action! – and it struck me that there is a problem between Heidegger (2010 [1927]) and Gadamer (2013 [1960]). Heidegger thinks interpretation is purely subjective, we can only read ourselves in text. Gadamer, however, thinks that works of others can expand our horizon.</p>

<p>The problem I find here with Gadamar is: how can we learn something we do not have first hand experience of? Can a child before it acquires language understand the consequences of touching the hot kettle on the table before it has experienced a – hopefully limitedly – terrifying event? I assume there exists gene encoded and therefore tacitly embodied intuitive knowledge – is the physiological reflex with which the child retracts its hand from the kettle such knowledge? Possibly, but that knowledge is fully embodied and has no conscience components, is seems to me. I do not see a convincing reason yet to assert gene transferred cognitive appreciations of any kind. However, the wailing that follows convinces me that there is now more conscience understanding resulting from the experience. Not of course as concrete as “kettles may be super hot, take care not to touch kettles without sufficient checking” – more some dim realization that linguistically later on might be put as “things can hurt”.</p>

<p>The child, mercifully is the malleable developing brain, will forget the concrete experience, but a part of understanding has been achieved.  The physiological experience extends to acquiring meaning as well. We cannot understand a meaning but before we have acquired it through an experience of some kind. The fact that information and understanding is able to flow between people by way of text (or speech) is therefore the ability to translate symbols and linguistic signs into a sort of imagined experience from which we then learn.</p>

<p>I think we should not confuse this with a reified semantics that is embedded in individual words or symbols or some linguistics connected to these. If I encounter a word that is fully and utterly new for me, I simply cannot understand it. I need it in a linguistic context of a sentence (and possibly a whole lot more sets of semantic signs, such as chapters and books) to help me have an imagined experience to understand – essentially performatively reconstruct – the meaning of that new word.</p>

<h3 id="references">References</h3>

<ul>
  <li>Hans-Georg Gadamar. 2013. <em>Turth and Method</em>. Translated by Joel Weinsheimer and Donald G. Marshall. First published 1975, 2nd edition 189, Revised edition 2004; Originally published 1960 (German). London, New York: Bloomsbury Academic.</li>
  <li>Heidegger, Martin. 2010. <em>Being and Time</em>. Translated by Joan Stambaugh. First published 1953, rev.with A foreword by Dennis J. Schmidt, Originally published in 1927 (German). Albany (NY): State University of New York Press.</li>
  <li>Zundert, Joris J. van. 2016. “Screwmeneutics and Hermenumericals: The Computationality of Hermeneutics.” In <em>A New Companion to Digital Humanities</em>, edited by Susan Scheibman, Ray Siemens, and John Unsworth, 331–347. Malden (US), Oxford (UK), etc.: John Wiley &amp; Sons, Ltd. <a href="http://onlinelibrary.wiley.com/doi/10.1002/9781118680605.ch23/summary">http://onlinelibrary.wiley.com/doi/10.1002/9781118680605.ch23/summary</a>.</li>
</ul>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[Thinking and reasoning about what interpretation exactly is, is an endless source of joyful wondering. I was just rereading my own Screwmeneutics (2016) – while preparing my thesis conclusion, there was no vanity in that action! – and it struck me that there is a problem between Heidegger (2010 [1927]) and Gadamer (2013 [1960]). Heidegger thinks interpretation is purely subjective, we can only read ourselves in text. Gadamer, however, thinks that works of others can expand our horizon.]]></summary></entry><entry><title type="html">Moving to Computational Humanities</title><link href="http://localhost:4000/moving-to-computational-humanities/" rel="alternate" type="text/html" title="Moving to Computational Humanities" /><published>2019-07-19T17:32:33+02:00</published><updated>2020-09-25T12:25:14+02:00</updated><id>http://localhost:4000/moving-to-computational-humanities</id><content type="html" xml:base="http://localhost:4000/moving-to-computational-humanities/"><![CDATA[<p><img style="float:left;padding-right:1em;width:250px;margin:0;border:none;" src="/assets/uploads/2019/folgert_max-354x1024.png" /><a href="https://www.karsdorp.io/">Folgert Karsdorp</a>, a colleague of mine over at the <a href="https://www.meertens.knaw.nl/cms/en/">Meertens Institute</a> doing impressively experimental work in the application of computing in the humanities voiced a recognizable unease yesterday by tweet (see on the left).<br /><br />It has become my understanding the past few months that the number of people active in the field of DH with such sympathies is on the rise. I have no idea if it is actually the case that computational intensive work is currently marginalized and pushed to the peripheries of Digital Humanities. But if so, that would raise serious concerns. It should be a serious signal for ADHO that excellent researchers feel the need to create new platforms like <a href="https://t.co/guCdu2s4Jj">CoHuRe</a>. Although ADHO’s wish to cater widely it may need to tend to how it also caters deeply.&lt;/p&gt;</p>

<p>In any case, I had some thoughts and tweeted them. And then <a href="https://www.maxkemman.nl/">Max Kemman</a> said it could be a blog post. So I turned the tweet stream into a blog post by applying a good old fashioned image map to it. Enjoy…</p>

<p><img style="width:600px;margin:0;border:none;" src="/assets/uploads/2019/twitterstream_20190719_1312.jpg" usemap="#image-map" /></p>

<!-- Image Map Generated by http://www.image-map.net/ -->

<map name="image-map">
    <area target="_blank" alt="" title="" href="https://t.co/ahK3NGPgSU" coords="94,157,562,279" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/CoHuRe1" coords="486,78,553,102" shape="rect" />
    <area target="_blank" alt="" title="" href="https://t.co/ahK3NGPgSU" coords="96,115,245,155" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/FolgertK" coords="441,457,378,438" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/CoHuRe1" coords="170,476,101,455" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/willardmccarty" coords="212,475,307,492" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/DH2019_NL" coords="438,803,521,819" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/CoHuRe1" coords="100,668,173,688" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/ADHOrg" coords="100,916,167,937" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/felwert" coords="100,63,249,78" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/MaxKemman" coords="102,323,277,338" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/MaxKemman" coords="216,2130,304,2147" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/MaxKemman" coords="267,2448,357,2465" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/tla" coords="102,2468,131,2485" shape="rect" />
    <area target="_blank" alt="" title="" href="https://t.co/YLmP2aJSfT" coords="373,2469,536,2487" shape="rect" />
    <area target="_blank" alt="" title="" href="https://t.co/YLmP2aJSfT" coords="103,2519,552,2616" shape="rect" />
    <area target="_blank" alt="" title="" href="https://t.co/lPLGw86Tzc" coords="102,1673,554,1792" shape="rect" />
    <area target="_blank" alt="" title="" href="https://t.co/lPLGw86Tzc" coords="381,1628,555,1644" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/FolgertK" coords="304,1611,371,1628" shape="rect" />
    <area target="_blank" alt="" title="" href="https://twitter.com/scott_bot" coords="226,936,152,954" shape="rect" />
    <area target="_blank" alt="I wrote another blog related to that" title="I wrote another blog related to that" href="http://jorisvanzundert.net/blogposts/why-we-should-think-about-a-domain-specific-computer-language-dsl-for-scholarship/" coords="156,1972,435,2003" shape="rect" />
    <area target="_blank" alt="Our DH2019 paper delves into realted and similar issues" title="Our DH2019 paper delves into realted and similar issues" href="https://dev.clariah.nl/files/dh2019/boa/0357.html" coords="97,2158,463,2215" shape="rect" />
    <area target="_blank" alt="Our DH2019 presentation delves into related and similar issues" title="Our DH2019 presentation delves into related and similar issues" href="http://jorisvanzundert.net/prezs/presentations/DH_2019/DH_2019.html#/Title" coords="96,2221,437,2277" shape="rect" />
</map>

<p><img style="float:left;padding-right:1em;width:250px;margin:0;border:none;" src="/assets/uploads/2019/Screen-Shot-2019-07-19-at-19.04.16-.png" alt="" class="wp-image-1252" /></p>

<p>The Twitter conversation forked and grew considerably in the hours that  followed those tweets. An intervention to mention was by <a href="http://eadh.org/fabio-ciotti">Fabio  Ciotti</a>—one of the Chairs of <a href="https://dh2019.adho.org/">ADHO’s DH2019 Conference</a> Programme Committee. According to Fabio the sympathies and feelings about  marginalization are at most anecdotal, the uses of DH for computer science negligible, and the merit of computational experimentation in the humanities doubtful. I guess Fabio felt under attack and shot from the hip. Follow up tweets seemed to be more inviting to a conversation. <br /><br />I am certainly not excluding the possibility that indeed the sense of  marginalization is something anecdotal rather than real. I asked if we can use ADHO’s database of reviews  to truly understand this issue. It  would be great if we can analyze this resource to get to the bottom. Curious to see where that goes.</p>

<p>Many  other conversations were sparked, which I guess is only good. For those  stories however, you’ll have to go hunt <a href="https://twitter.com/true_mxp/status/1152254310102777856">Twitter</a> itself.</p>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[Folgert Karsdorp, a colleague of mine over at the Meertens Institute doing impressively experimental work in the application of computing in the humanities voiced a recognizable unease yesterday by tweet (see on the left).It has become my understanding the past few months that the number of people active in the field of DH with such sympathies is on the rise. I have no idea if it is actually the case that computational intensive work is currently marginalized and pushed to the peripheries of Digital Humanities. But if so, that would raise serious concerns. It should be a serious signal for ADHO that excellent researchers feel the need to create new platforms like CoHuRe. Although ADHO’s wish to cater widely it may need to tend to how it also caters deeply.&lt;/p&gt;]]></summary></entry><entry><title type="html">Why We Should Think About a Domain Specific Computer Language (DSL) for Scholarship</title><link href="http://localhost:4000/why-we-should-think-about-a-domain-specific-computer-language-dsl-for-scholarship/" rel="alternate" type="text/html" title="Why We Should Think About a Domain Specific Computer Language (DSL) for Scholarship" /><published>2018-10-30T15:15:56+01:00</published><updated>2019-05-14T16:34:50+02:00</updated><id>http://localhost:4000/why-we-should-think-about-a-domain-specific-computer-language-dsl-for-scholarship</id><content type="html" xml:base="http://localhost:4000/why-we-should-think-about-a-domain-specific-computer-language-dsl-for-scholarship/"><![CDATA[<h2 id="introduction">Introduction</h2>
<p>This is the text of a paper I presented during the conference “Digital Hermeneutics in History: Theory and Practice”, organized by the C<sup>2</sup>DH of Luxembourg University on 25 and 26 October 2018. I have been toying with the idea for a Domain Specific Language for textual scholarship for over a decade. Manfred Thaller—not aware as far as I know of my idea—has propelled it out of its dull momentum again for me by a <a href="https://ivorytower.hypotheses.org/56">wonderful blogpost</a> that he published in April (Thaller 2018). Manfred’s propositions seem to converge to a great extent with mine, and this text is greatly indebted to his blogpost. I intend to develop my preliminary argument here into an article describing a more concrete proposal for a DSL. This post is therefore a heartfelt Request For Comments (RFC) in the hope that interested people from all sides of the involved spectrum—computer scientists, software engineers, textual scholars and Science &amp; Technology Studies researchers—will want to enter a into a dialogue on the subject</p>

<h2 id="1-the-watershed-anecdote">1. The Watershed Anecdote</h2>
<p>I have been pitching the idea of a humanities specific computer language for years now. The pitch has become an almost perfect tool to separate scholars in the humanities on the one hand from software engineers and digital humanists on the other hand.</p>

<p>Scholars immediately love the idea of a computer language that is modelled after their knowledge domain and uses verbs and lexical items they can actually understand and apply. But engineers and digital humanists at best will roll their eyes when I mention the idea, and in some worse cases they indeed got very annoyed at the suggestion.</p>

<p>I think this watershed points us to a problem that we need to address: scholars feel out of their depths with current main stream general purpose computer languages. It seems however that the persons that could address it (software engineers, computer scientists, and digital humanists) are not too interested in pursuing a solution.</p>

<p>This problem has many roots in the history of philosophy and science, but here I want to cast it specifically as a problem of digital hermeneutics.</p>

<h2 id="2-two-main-problems">2. Two Main Problems</h2>
<h3 id="21-problem-one-most-scholars-are-computationally-illiterate-and-digital-hermeneutics-requires-code-literacy">2.1. Problem One: Most Scholars are Computationally Illiterate and Digital Hermeneutics Requires Code Literacy</h3>
<p>Hermeneutics is the theory and the methodology of interpretation. For more than two millennia it has been the main critical tool of humanists to examine texts and other sources of information. Its importance for a large part stems from the pivotal task of interpreting the Holy Scriptures during medieval and early modern times. An interesting aspect of hermeneutics is that method and object mostly coincide: in the far majority of cases text is interpreted and criticized by creating more text. In a culture where text is the main medium of knowledge creation and exchange this makes sense.</p>

<p>It is doubtful however if text is indeed still the linchpin medium of current human culture. Arguably it is more precise to say that nowadays digital information is the main medium of knowledge creation and exchange. Therefore scholarship should in part mean being able to interpret digital information. As pointed out however, virtually no scholar has the skills to do so.</p>

<p>The arguments against my proposals for a DSL for scholarship usually run largely along the line that one does not need to understand how a printing press works to read print, or that one does not need to build a printing press, or even operate it, to do so. Similarly therefore—goes the reasoning—one does not need to understand software code to appreciate the message it conveys in the form of text and images in a Graphical User Interface (GUI). Besides, the opponents reason, scholars are interested in the underlying processes and the methods and not in a particular piece of code which is just a particular expression of those analytic processes.</p>

<p>Such reasoning is dangerously flawed and creates a very real problem of digital hermeneutics, because it denies a primary site of interpretation: the software code that underpins all digital information.</p>

<p>Let me explain this further by first explaining what happens when we interpret.</p>

<h4 id="211-what-is-interpretation">2.1.1. What is Interpretation?</h4>
<p>It is a little ridicule to suggest that yards of shelf space including that for Gadamer’s tome “Wahrheit und Methode” can be summarized in one short and neatly universal fitting definition of ‘interpretation’. However, for the argument here it is sufficient to use a slightly circular working definition of “interpretation” as the applied or practical form of hermeneutics. Umberto Eco in <em>The Theory of the Sign and the Role of the Reader</em> (1981) describes text essentially as a stream of reading instructions. Interpretation, or meaning creation, is then the iterative and reflexive process based on these reading instructions. Eco opposes this process of meaning creation to the idea of meaning as a linguistic identity. What he mean—or at least how I take it—is that there is no identity between sign and meaning. A sign ‘a’ is not locked to meaning ‘b’ (a ≡ b), rather a sign is ‘adorned’ with meaning during the reflexive process of interpretation which is influenced by both the linguistic context in the text and the semantic context in the mind. This process and reflexivity create what Charles Sanders Peirce called ‘infinite semiosis’, which in effect is the ability of signs to infinitely evoke new meaning.</p>

<p>From this fast paced comprehensive stride through theory we derive the useful observations that interpretation is the attribution of a particular meaning to information and that meaning is an emergent property of processing information. The salient point being that interpretation—or meaning attribution—is a by all means a <em>process</em>.</p>

<h4 id="212-why-the-usual-metaphors-are-flawed">2.1.2. Why the Usual Metaphors are Flawed</h4>
<p>My argument is that because information processing is essential for interpretation the metaphors usually set in opposition to a Domain Specific Language for the humanities are flawed: they create a false image of processing information without interpretation—as if code solely results in neutral transformations of information. The metaphors pitched against a DSL for scholarship (or against coding by humanists in general) picture code as a thing instead as a process. The two most used are ‘car’ and ‘microscope’.</p>

<p>The car metaphors comes in variants, but usually the engine is compared to the computer, the gasoline to the software, and the driver obviously to the user. The metaphor of the car suggests that a computer is mechanical immutable and that software is a homogeneous neutral force. It also suggest the user is unproblematic in control of both. The microscope metaphor has probably been popularized through Ian Hacking’s widespread philosophy of science piece “Do We See Through a Microscope?” (1981), although he did not use it as a metaphor for computer or software. The metaphor of the computer or software as a microscope suggests again that code is neutral by comparing it how a microscope is physically and materially neutral to the passing of light which is governed by a mathematically neutral theory of optics and tested with known visual materials.</p>

<p>Of course these metaphors by themselves are already completely flawed. Software is everything but a homogeneous substance as any software engineer will tell you. Software is understood mostly as a technical and impartial tool, but it is also or even primarily a culturally situated human made artefact.</p>

<p>But the usual metaphors are especially pernicious in what they want to suggest: that one does not need to know how these machines work to operate them and interpret their results. Of course you need to know something about the working of a microscope, even if it is just the rudimentary knowledge that it “magnifies”. If you would not have that knowledge you would take a (to you) completely harmless animal the size of a speck of dust as a three foot tall terrifying creature. Especially if we look at the concrete effect of a microscope we realize that what it does is very much a strong <em>interpretative move</em>. It turns a harmless speck of dust into a scary threat. It most literally changes one’s perspective.</p>

<p><a href="/assets/uploads/2018/micromonster_blog.jpg"><img style="margin-bottom:0.5em;" src="/assets/uploads/2018/micromonster_blog.jpg" alt="Fig. 1: Diving beetle larva, pretty harmless in real life at 2cms, kind of scary at the right magnification. (Image reproduced from https://cosmosmagazine.com/biology/micro-monsters-up-close-and-personal, credit: EYE OF SCIENCE / SCIENCE PHOTO LIBRARY / GETTY IMAGES.)" /></a><br />
<em>Fig. 1: Springtail, pretty harmless in real life at 6mm, kind of scary at the right magnification. (Image reproduced and cropped from https://cosmosmagazine.com/biology/micro-monsters-up-close-and-personal, credit: Eye Of Science / Science Photo Library / Getty Images.)</em></p>

<h4 id="213-when-is-hermeneutics-digital-hermeneutics">2.1.3. When is Hermeneutics <em>Digital</em> Hermeneutics</h4>
<p>Is there anything inherent methodologically digital in analyzing a number of historic paintings from super high resolution digital images projected on a screen? If the hermeneutics applied are identical to the hermeneutic method applied in conventional art history, I would argue that this hermeneutics is indeed <em>not</em> in any way inherently methodological digital.</p>

<p>The processual dimension of hermeneutics requires that we take the full process of interpretation into account if we are to speak of true hermeneutics—i.e. meticulous methodic interpretation. The digital dimension of digital hermeneutics requires us to take into account the specifically digital aspects of that process.</p>

<p>Digital hermeneutics without taking into account the digital processual dimension compares to wanting to interpret a football match by the resulting score of 3 to nil alone. Methodic hermeneutics is interpreting the football match by actual watching it in process. Interpreting a football match by the final score alone can be a valid form of interpretation when it fits the purpose, e.g. to gather score statistics over a multitude of matches. But reading the match in a hermeneutic way requires a different approach, that of watching and experiencing it. Similarly digital hermeneutics requires an equally intimate experience of code that is involved in the process of interpretation.</p>

<p>Thus a truly digital hermeneutics involves a hermeneutics of the specifically digital parts of the interpretation process. But this is not to say that all digital parts of such a process require close reading.</p>

<p>Hinsen (2017) makes a useful distinction that divides scientific software code into four layers. A first layer of general infrastructural software (operating systems, compilers, generic tools like text processors). A second layer comprises scientific software of specifically science aimed applications and libraries (SPSS, Stata, Gnuplot). A third layer comprises disciplinary software and libraries (Stylo, NLTK). Finally there is bespoke code: specific one off tailored software, used mostly transiently in a single project.</p>

<p>A ‘black boxing’ hermeneutics that is primarily interested in interpreting the results of a tool chain that is largely made up of thoroughly evaluated and and continuously tested generic applications and libraries may well forego on including that digital code as part of the hermeneutic process, but taking for granted the interpretational dimension of bespoke code would be methodological hazardous.</p>

<p>Especially bespoke code is a mechanism of interpretation that a user needs to closely read and understand if it is part of <em>digital</em> hermeneutic interpretation or evaluation.</p>

<p>Mutatis mutandis what goes for data—namely that the more heterogeneous and situated one’s data is the more hermeneutic interrogation it requires—goes for code as well: the more situational the code used the more critically a hermeneutic process needs to examine it and take it into account for any interpretation.</p>

<h4 id="214-how-code-contributes-to-interpretation">2.1.4. How Code Contributes to Interpretation</h4>
<p>Code can be both transformative and generative. In its transformative guise code shifts the shape of data or information. In its generative mode it adds to information to create new or augmented information. Especially of code in its transformative guise many people hold that it is neutral, that it does not affect the essential meaning of the information it is processing. But in fact it is of course exactly doing that in veritably all cases. Take as an example a simple mapping transform where a multiplier is applied to a sinus.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>o = (0.0..10.0).step(0.1).map{ |x| Math.sin(x) }
m = (0.0..10.0).step(0.1).map{ |x| 20 * Math.sin(x) }
</code></pre></div></div>

<p><a href="/assets/uploads/2018/sin_combined.png"><img style="margin-bottom:0.5em;" src="/assets/uploads/2018/sin_combined.png" alt="A mapping from sin(x) to 20*sin(x)" /></a> 
<em>Fig. 2.: Depiction of how ‘o’ and ‘m’ look when plotted.[/caption]</em></p>

<p>Just like the microscope this chances the perspective dramatically even if the original information is the same. Such a multiplier effect is—even if it is unintentionally—also an interpretative move: it privileges certain information, subdues other signals, foregrounds a certain perspective, and primes a certain interpretation.</p>

<p>In its generative mode these interpretative moves become only more pronounced. When code combines or infers information it becomes almost impossible to speak of such operations as non-interpretative. If, for instance, one is aggregating frequency information of a large corpus of literary texts and adding a logistic regression revealing a shift in elite literary word use in the long 19th century, like Ted Underwood (2015) it is rather rich not to speak of an explicit interpretative move through the code. In the sense that I have written earlier of the potential ‘deferred agency’ of code, one could call this deferred analytics or indeed deferred interpretation. The human interpreter delegates the interpretative move to the code.</p>

<h4 id="215-back-to-interpretation-or-the-necessity-of-literacy">2.1.5. Back to Interpretation, or: The necessity of Literacy</h4>
<p>My reasoning so far adds up to the conclusion that interpretation is spread throughout the mechanisms of code, code writing, and code application. Interpretation is fully intertwined with research method and process, and thus also with code and computation that are integral parts of that process.</p>

<p>Now please interpret this…</p>

<p><a href="/assets/uploads/2018/udhr_burmese.gif"><img style="margin-bottom:0.5em;" src="/assets/uploads/2018/udhr_burmese.gif" alt="Fig. 3.: Sample of Burmese script, reproduced from Simon Ager (2018) &quot;Omniglot - writing systems and languages of the world&quot;, available at https://www.omniglot.com/writing/burmese.htm (accessed 30 October 2018)." /></a><br />
<em>Fig. 3.: Sample of Burmese script, reproduced from Simon Ager (2018) “Omniglot - writing systems and languages of the world”, available at https://www.omniglot.com/writing/burmese.htm (accessed 30 October 2018).</em></p>

<p>Foregoing on the small chance that you actually read Burmese—and apologizing to the reader that I am assuming an audience for this article of mostly ‘cultural westerners’—this should suffice to show that there is nothing to interpret here. This is because you, reader, are illiterate in Burmese.</p>

<p>Now please interpret the difference between the following two simple two dimensional arrays.</p>

<p><a href="/assets/uploads/2018/arr_frac_num.png"><img style="margin-bottom:0.5em;" src="/assets/uploads/2018/arr_frac_num.png" /></a></p>

<p><a href="/assets/uploads/2018/arr_rand_num.png"><img style="margin-bottom:0.5em;" src="/assets/uploads/2018/arr_rand_num.png" alt="Fig. 4.: Two simple comparable two dimensional matrices." /></a><br />
<em>Fig. 4.: Two simple comparable two dimensional matrices.</em></p>

<p>Of course we experience exactly the same thing here. Apart from maybe some ‘savants’ we are unable to see what this data might be. For interpretation we are dependent on our code microscope to augment our interpretation, which might interpret these numbers as pictures.</p>

<p><a href="/assets/uploads/2018/arr_frac.png"><img style="margin-bottom:0.5em;" src="/assets/uploads/2018/arr_frac.png" /></a></p>

<p><a href="/assets/uploads/2018/arr_rand.png"><img style="margin-bottom:0.5em;" src="/assets/uploads/2018/arr_rand.png" alt="Fig. 5.: The matrices of figure 3 represented as grid surfaces." /></a><br />
<em>Fig. 5.: The matrices of figure 3 represented as grid surfaces.</em></p>

<p>To help us see the rather striking difference between these two sets of numbers we use code that adorns the sets with a first interpretation. Indeed like a microscope the code functions as an extension of our eyes turning these numbers more interpretable.</p>

<p>In the cases above we are almost fully illiterate: we cannot read or interpret either the data, the code, nor the interpretation the code creates.</p>

<p>Many before me have argued that if it is unknown what data means or pertains to then we cannot reliably interpret the data (Borgman 2015, Gitelman 2013). Many have likewise argued that if we cannot exactly know how a method transforms the data, then at best interpreting the results of such transformations becomes cargo culture, and at worst it results in nonsense. Because code is, as argued, a process with a similar interpretational aspect, we can only ignore it at the peril of creating illusory interpretations.</p>

<p>This is all to say that you cannot perform valid interpretation if you lack the ability to evaluate the interpreting method. In the case of code this means that at some level code literacy is needed to evaluate the validity of any interpretation derived in a process that included the application of that specific code.</p>

<h3 id="22-problem-two-current-general-purpose-computer-languages-suck-from-the-hermeneutic-point-of-view">2.2 Problem Two: Current General Purpose Computer Languages Suck From the Hermeneutic Point of View</h3>
<h4 id="221-coding-is-hard-and-engineer-oriented">2.2.1. Coding is Hard and Engineer Oriented</h4>
<p>Although from a digital hermeneutics perspective code should be examined as part of the interpretation process, current general purpose computer languages try to do their utmost to keep scholars out of the process.</p>

<p>Having programmed since the early 1980s it surprises me still that so called high level languages are still so low level in many respects. A thing as simple as reading a file is surrounded by a baffling amount of engineering clutter. The shortest one can do in Ruby is:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>File.open( "my_data.txt", "r" ) { |file| file.read }
</code></pre></div></div>

<p>The Python is similar:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>with open( 'pagehead.htm', 'r' ) as file:
output = file.read()
</code></pre></div></div>

<p>One wonders what is wrong with or impossible about:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Use "some_file"
</code></pre></div></div>

<p>The problem is that current main stream computer languages are fully enveloped in an object oriented philosophy that puts the objects, structures, and operations of file system and data sources first and centre. They operate primarily at a level just up from the operating system, but with objects and abstractions very much below the level of any content domain. Quite logically of course, as they are indeed general purpose languages geared towards software engineers. They are languages made by engineers for engineers. They offer the user a code execution environment with the technical objects familiar to the coder: input and output access to the file system, some primitives like string, integer, and array, a few flow control and boolean operators. The idea is that from these primitives coders build a conceptual object oriented perspective of the problem domain. Some libraries indeed go a long way in that direction, although most are statistics or NLP oriented (NumPy, NLTK).</p>

<p>The result is that a scholar who wants to manipulate domain information finds him or herself wasting most of the time in dealing with the scaffolding stuff of system calls, pipes, and data structures rather than with actual analytic interpretative code.</p>

<p>This problem is aggravated by a vicious circle of software development and architectural bloat that has continued for over twenty years. There is a much longer story to tell about this (Prokopov 2018), but suffice to say that once it was plenty enough to know 13 HTML tags to be a web developer, while now it takes knowledge of HTML, CSS, JavaScript, NodeJS, Angular, Grunt, etc. etc. Next to a single language, software engineers nowadays often will use development frameworks, web frameworks, testing frameworks, integration frameworks, build frameworks, virtualization frameworks, and deployment frameworks to make their work smooth sailing.</p>

<h4 id="222-information-vs-data">2.2.2. Information vs Data</h4>
<p>Technologists—computer scientists, software engineers, digital humanists—generally think of computing as the modelling, transformation, and analysis of <em>data</em>. However, scholars with a hermeneutic propensity are most of the time not interested in computation over data but in reasoning over information. Information usually is defined as data in context: “1 degree Celsius” is data, “1 degree Celsius on 5 July 1995 in Amsterdam” is information. The current patterns of code development are primarily focused on structuring and transforming data to visualize it in ways that end users may interpret it in a graphical user interface. This approach to computing drives the computational and the hermeneutic paradigms rather apart than towards each other. (For a more in depth treatment of this aspect turn to Thaller 2018.)</p>

<h4 id="223-algorithm-vs-reasoning">2.2.3. Algorithm vs Reasoning</h4>
<p>By virtue of serving to almost all the needs of industry veritably all general purpose computer languages are first order logic languages—also called predicate logic languages. The ramification is that general purpose computer languages are hermeneutically poor because out of the box they only support boolean reasoning. Hermeneutics is not based on boolean comparisons, but on abductive reasoning: it tries to find the most plausible argument that combines or explains sparse and heterogeneous data.</p>

<p>The absolutism of boolean reasoning implies that variables in general purpose computer languages can only hold one value:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>a = b
</code></pre></div></div>

<p>Abductive reasoning however requires a lot of leeway in the value of any variable, because it allows for ambiguity, uncertainty, possibility rather than probability. Thus in abductive reasoning</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>a = b &amp; a = c
</code></pre></div></div>

<p>is an assignment, not a comparison.</p>

<p>A lot of these characteristics of abductive reasoning can actually be supported through probabilistic computing. Indeed general purpose languages can come a long way in supporting probabilistic reasoning. However, this has to be done by creating several abstraction layers in such a language that together then make up a probabilistic reasoning engine. None of them is probabilistic at the ‘bare metal’. But even if they were: probabilistic reasoning is still not exactly abductive reasoning. Probabilistic reasoning considers a statement true because it is ‘probable’, and it is probable because there are enough prior examples of the statement being true (e.g. all ripe bananas are yellow). Abductive reasoning however, even allows for leaps of imagination that go beyond ‘plausible’ and have little or no support in prior examples. In abductive reasoning facts may be ‘imaginable’, for instance, rather than ‘probable’ or ‘true’.</p>

<p>These are probably the most fundamental differences between current general purpose computer languages and a more hermeneutically inspired computer language.</p>

<h3 id="3-problems-1-and-2-lead-to-the-central-problem-the-renaissance-man-deadlock">3. Problems 1 and 2 Lead to the Central Problem: the Renaissance Man Deadlock</h3>
<p>I have demonstrated that there is a real need for scholars to engage with code and to engage with code in a hermeneutic fashion. Current general purpose computer languages however because of their low level, data oriented, first order logic nature resist such hermeneutic engagement. These languages are thus not very inviting or attractive for scholars to learn and to apply.</p>

<p>Most certainly general purpose computer languages can be stretched to serve as more hermeneutically inclined tools, e.g. along probabilistic approaches. The problem however, is that to do so requires a level of code literacy that is even beyond the coding fluency that is needed to just understand general purpose computer languages in their guise as first order computing languages.</p>

<p>Thus to have scholars work around the fact that current general purpose computer languages are pretty horrible digital hermeneutic instruments, you need to ‘upgrade’ these scholars to a level of computer language proficiency that is not even average in the software engineering domain.</p>

<p>But a person can be only so much expert in so many fields. You cannot turn humanities scholars into expert coders and have them being expert humanists as well. It is nigh impossible to be both an expert coder and simultaneously be an expert in a humanities topic sufficiently to be productive at an academic level. This is what I mean by the Renaissance Man Deadlock: if you try to be expert in both directions, you will suck at both. Excelling at both makes you a kind of scientific unobtainium (Unobtainium 2018).</p>

<h3 id="4-towards-a-solution-if-the-mountain-will-not-come-to-mohammed-mohammed-will-go-to-the-mountain">4. Towards a Solution: If the Mountain Will Not Come to Mohammed, Mohammed Will Go to the Mountain.</h3>
<p>Or in other words: we need better computer languages that are more geared towards hermeneutics.</p>

<h4 id="41-what-might-a-dsl-for-scholarship-look-like">4.1. What Might a DSL for Scholarship Look Like?</h4>
<p>First of all such a language would be far more high level than anything we are used to so far. It would certainly abstract away completely from the low level file system and piping calls that are now often needed to accomplish anything useful but that require the bulk of the programming time. So it indeed would implement something like <code>Use "my_data.json"</code>.</p>

<p>This suggestion itself may send IT architects screaming for all the security holes and error prone assumptions about the underlying filesystem such a language would result in. But the salient point is that computer scientists and software engineers always face in the direction of their platforms. What I invite them to do is to instead focus a while on the question “What is the most convenient way for the scholarly users to reach for their data?”</p>

<p>A hermeneutic computing language’s unit of processing is information not data. Before claiming impossibilities, engineers and computer scientists should note that it is actually normal for a unit of processing to be at a higher level of bits and bytes. Data is usually already a complex structured digital object, but the challenge is not to leave it at that.</p>

<p>Manfred Thaller (2018) in this respect suggests to implement Keith Devlin’s idea of <em>infons</em> “for seamless usage in main stream programming languages”. An infon could be formally represented as:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>   « P, a1, a2, … t, p, 0.1 »
</code></pre></div></div>

<p>Meaning that certain objects (a1, a2, etc.) at a certain time (t) and place (p) have a determined relationship P with a probability of 0.1.</p>

<p>My quibble here would be not with the infons (I think they make sense as an atomic unit of hermeneutic information) but with “seamless usage in main stream programming languages”. As stated I think main stream programming languages are not suited for scholarship because their idiom remains at too low a level of abstraction. Infons should indeed better be implemented as part of a hermeneutics geared domain specific computer language.</p>

<p>A hermeneutic computing language or a DSL for scholarship should be as polyglot as possible. Coding often means spending huge amounts of time and effort piping data from one language to another (e.g. feeding data into Gnuplot from Ruby), this should be far more convenient:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>def plot_grid&lt;
   grid_data = @data_points
   R.set style line 2 lc rgb '#5e9c36' pt 6 ps 1 lt 1 lw 2
   R.set dgrid3d 50,50 gauss 4
   R.splot grid_data u 1:2:3 w lines
puts "done"&lt;/code&gt;
</code></pre></div></div>

<p>This DSL I would like to see supports some form of ‘beyond boolean’ logic. For instance, the statement</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>  a = b &amp; a = c
</code></pre></div></div>

<p>in such a language would be a value assignment instead of a boolean comparison (which would conventionally be denoted by using <code>"a == b &amp;&amp; a == c"</code>). The assignment itself would mean that <code>a might be b</code> or <code>a might be c</code>. This implies that such a language would allow for simultaneous computation of equivalent execution paths to compare hypotheses cast by such assignments.</p>

<p>A language like this might be probabilistic, but must in any case be ‘fuzzy’ to support ambiguity, imprecision, and uncertainty—because in “ca. 1564” the amount that circa signifies varies by context.</p>

<p>Finally a hermeneutic language would implement more flow control operators than just the conditional ‘if’. It would, for instance, implement ‘therefore’ and ‘because’ like operators, which are equivalent to ‘if’ but manipulate the order and dependency structure of argument.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>I will go hungry if I do not eat
I will go hungry because I do not eat
I do not eat therefore I go hungry
</code></pre></div></div>

<h3 id="5-conclusion">5. Conclusion</h3>
<p>Main stream general purpose computer languages are seriously flawed as hermeneutic computing languages and scholars are seriously challenged in acquiring fluency in such languages because they cannot be expert in two fields simultaneously.</p>

<p>Truly <em>digital</em> hermeneutics however require a deep and intimate engagement with the interpretative aspects of code, because interpretation is a process of which code can be, and often is, an integral part.</p>

<p>It is therefore that we should at least think about a domain specific computing language that will allow scholars to enter the hermeneutic space that code creates. Humanities critical role could be seriously impaired if scholars were not allowed entrance to the sites of interpretation that computer languages have created.</p>

<p>Indicative properties and broad outlines of a domain specific language for scholarship can be given. Also prior work in computer science exists that indicates that such a language falls within the bounds of the possible.</p>

<p>The ultimate challenge is to build a hermeneutics oriented computer language from the ground up. However, as first practice and baby steps higher level languages like Python and Ruby should offer plenty of sandbox space for exercises towards such an eventual new language.</p>

<p>To prevent it becoming some idiosyncratic idiolect, designing a language like the envisioned domain specific computer language for scholarship should not be some individual endeavour. Rather it should be an interdisciplinary and community driven development, which could well be modelled after or be part of the ‘Shared Task in the Digital Humanities’ (Reiter 2016).</p>

<p>–JZ_20181019_1551</p>

<h3 id="references">References</h3>

<ul>
  <li>Borgman, C. L. (2015) <em>Big Data, Little Data, No Data: Scholarship in the Networked World</em>. Cambridge Mass.: MIT Press.Eco, U. (1981) ‘The Theory of Signs and the Role of the Reader’, <em>The Bulletin of the Midwest Modern Language Association</em>, 14(1), pp. 35–45. doi: 10.2307/1314865.</li>
  <li>Gitelman, L. (ed.) (2013) <em>“Raw Data” Is an Oxymoron</em>. Cambridge (MA), USA: The MIT Press.</li>
  <li>Hacking, I. (1981) ‘Do We See Through a Microscope?’, <em>Pacific Philosophical Quarterly</em>, 62(4), pp. 305–322.</li>
  <li>Hinsen, K. (2017) ‘Sustainable software and reproducible research: dealing with software collapse’, Konrad Hinsen’s Blog, 13 January. Available at: <a href="http://blog.khinsen.net/posts/2017/01/13/sustainable-software-and-reproducible-research-dealing-with-software-collapse/">http://blog.khinsen.net/posts/2017/01/13/sustainable-software-and-reproducible-research-dealing-with-software-collapse/</a> (Accessed: 21 April 2017).</li>
  <li>Prokopov, N. (2018) Software disenchantment, tonsky.me. Available at: <a href="http://tonsky.me/blog/disenchantment/">http://tonsky.me/blog/disenchantment/</a> (Accessed: 25 October 2018).</li>
  <li>Reiter, N. (2016) An Initiative for Shared Tasks in the Digital Humanities, GitHub. Available at: <a href="https://github.com/SharedTasksInTheDH">https://github.com/SharedTasksInTheDH</a> (Accessed: 13 July 2018).</li>
  <li>Thaller, M. (2018) ‘On Information in Historical Sources’, A Digital Ivory Tower, 24 April. Available at: <a href="https://ivorytower.hypotheses.org/56#more-56">https://ivorytower.hypotheses.org/56</a> (Accessed: 16 June 2018).</li>
  <li>Underwood, T. and Sellers, J. (2015) How Quickly Do Literary Standards Change?, Figshare. Available at: <a href="http://figshare.com/articles/How_Quickly_Do_Literary_Standards_Change_/1418394">http://figshare.com/articles/How_Quickly_Do_Literary_Standards_Change_/1418394</a> (Accessed: 16 December 2015).</li>
  <li>Unobtainium (2018) tvtropes. Available at: <a href="https://tvtropes.org/pmwiki/pmwiki.php/Main/Unobtainium">https://tvtropes.org/pmwiki/pmwiki.php/Main/Unobtainium</a> (Accessed: 25 October 2018).</li>
</ul>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[Introduction This is the text of a paper I presented during the conference “Digital Hermeneutics in History: Theory and Practice”, organized by the C2DH of Luxembourg University on 25 and 26 October 2018. I have been toying with the idea for a Domain Specific Language for textual scholarship for over a decade. Manfred Thaller—not aware as far as I know of my idea—has propelled it out of its dull momentum again for me by a wonderful blogpost that he published in April (Thaller 2018). Manfred’s propositions seem to converge to a great extent with mine, and this text is greatly indebted to his blogpost. I intend to develop my preliminary argument here into an article describing a more concrete proposal for a DSL. This post is therefore a heartfelt Request For Comments (RFC) in the hope that interested people from all sides of the involved spectrum—computer scientists, software engineers, textual scholars and Science &amp; Technology Studies researchers—will want to enter a into a dialogue on the subject]]></summary></entry><entry><title type="html">Being a Critical Journalist in Digital Times</title><link href="http://localhost:4000/being-a-critical-journalist-in-digital-times/" rel="alternate" type="text/html" title="Being a Critical Journalist in Digital Times" /><published>2017-07-10T21:25:51+02:00</published><updated>2018-02-09T14:06:59+01:00</updated><id>http://localhost:4000/being-a-critical-journalist-in-digital-times</id><content type="html" xml:base="http://localhost:4000/being-a-critical-journalist-in-digital-times/"><![CDATA[<p><a href="http://networkcultures.org/geert/">Geert Lovink</a> who does wonderful work at the <a href="http://networkcultures.org/">Institute of Network Cultures</a> yesterday tweeted a <a href="https://twitter.com/glovink/status/883934260268343296">cry of horror</a> on finding out via <a href="https://techcrunch.com/2017/07/08/google-is-funding-the-creation-of-software-that-writes-local-news-stories/">TechCrunch</a> that Google is funding the development of software that writes local news stories. <a href="https://www.google.nl/search?q=Reporters+And+Data+And+Robots+Digital+News+Initiative">Media the world</a> over have parroted the same news which seems largely based on a <a href="https://www.pressassociation.com/company-news/pa-awarded-e706000-grant-google-fund-local-news-automation-service-collaboration-urbs-media/">press release</a> from the <a href="https://www.pressassociation.com/">UK Press Association</a>. The parroting in itself is a indicator of the dire situation in journalism where uncritically posting press releases has become a stand in for actual in depth and well researched coverage. Those who at least attempted a stab at a perspective mostly seem to have stuck to the hackneyed criticism that Google funds the development of <a href="http://globalnews.ca/news/3580507/press-association-google-grant-robot-journalism/">robot journalists</a> that will <a href="http://mashable.com/2017/07/07/google-ai-journalism-funding-europe/#UWKuyV5A.sqs?">put human journalists out of a job</a>: “Journalists, look out: Google is funding the rise of the AI news machine”.</p>

<p>In reality things probably move both faster and slower. Let me explain that.</p>

<p>Having been in the midst of <a href="https://www.nrc.nl/nieuws/2017/03/23/met-enige-aarzeling-omarmen-de-uitgevers-big-data-als-hun-redding-7528498-a1551555">my own little media tempest in a teacup</a> I tend to think that press releases and the follow up roar in diverse media are as much ‘alternative fact’ as they are not. My case involved also a director of a publishing house announcing that we as researchers were ragingly enthusiast about the effectiveness with which we are able to predict best sellers based on deep learning methods. I think I—which is “the researchers” in the story—said results were encouraging or some such, how that turned to “ragingly enthusiast” remains a mystery to me too.</p>

<p>What I did in any case was not exactly rocket science in the realm of machine learning. Using the well established open source Python libraries <a href="http://deeplearning.net/software/theano/">Theano</a> and <a href="https://keras.io/">Keras</a> I build a straight forward neural network that I fed the 250,000 or so features counting <a href="https://en.wikipedia.org/wiki/Tf%E2%80%93idf">Tf·idf matrix</a> that was derived from a set of 200 Dutch published novels of which sales numbers were known. We were then able to predict the ‘selling capability’ of unseen novels of the same publisher for which sales numbers were known with some 80% accuracy. In more plain English: applying meanwhile middle of the road machine learning techniques we can predict if novels will sell or not and eight out of ten times we will be correct.</p>

<p>Machine learning techniques such as deep learning using neural networks can be extremely sensitive to patterns in large data sets, patterns that are too distributed throughout the data for humans to be really able to pick up on them. Given enough training data such technologies infer models, or sets of features, that will be very good at telling you to what categories unseen examples belong. If they belong to the category of best sellers or non sellers for instance. In our case the model picked up on the words that are common to the best sellers of recent years and was able to find matches of such word use in novels it had not been trained on. As a meticulous and time unlimited clerk it compared the scores on some 250k variables per novel, averaged the scores, and compared them to those of successful selling novels. Not exactly rocket science really, just an immense amount of work impossible to pull off in feasible time with mere human capacity. For a well programmed algorithm though, a work of mere minutes.</p>

<p>Classifiers, as such algorithms are also called, are already crunching real world data all the time. This is what I mean with the ‘faster’ part. In many respects what AP is going to try to do, is not that new at all. Stock exchange predictions, flight fuel consumption patterns, internet store customer interest, the likelihood a person on Facebook will lean to a certain political conviction, and so forth: all have been measured and predicted by similar algorithms for a small decade now at least.</p>

<p>The information that citizens experience is more often than not tailored to their needs already by such algorithms. Although many point this out all the time (for instance <a href="https://www.theguardian.com/technology/2013/mar/09/evgeny-morozov-technology-solutionism-interview">here</a>, or <a href="https://www.theoryculturesociety.org/interview-with-david-berry-on-digital-power-and-critical-theory/">here</a>, or <a href="http://rhizome.org/editorial/2013/jul/10/lev-manovich-interview/">here</a>, or <a href="http://www.newyorker.com/magazine/2015/11/23/doomsday-invention-artificial-intelligence-nick-bostrom">here</a>) it is somehow still a huge surprise when sometimes the processes that these algorithms support break out of their otherwise mostly covert and invisible existence. To those that are aware of how widespread the application of these algorithms are, rather this very surprise is surprising: you are living in an highly automated information world already. Long time. Better get used to it. Or at least get aware of it. And yes, this also happens in news story generation already, as <a href="https://techcrunch.com/2017/07/08/google-is-funding-the-creation-of-software-that-writes-local-news-stories/">TechCrunch</a> did not fail to <a href="https://www.google.nl/search?client=safari&amp;rls=en&amp;q=apple+%22this+story+was+generated+by+Automated+Insights%22">point out</a>. <a href="https://automatedinsights.com/">Automated Insights</a> provides this type of services on impressive industry scale.</p>

<p>The ‘slower’ bit is that press releases like the one from AP systematically over claim. Although I have to admit this is conjecture, it is unlikely that Google’s 700kEuro+ funding will lead AP and Urbs Media to eradicate local news gathering and publishing. First of all this does not seem their aim, but moreover the quality and effects of “a new service” crunching out “up to 30,000 localised stories each month from open data sets” remain to be seen. What are these open data sets, and what will these news stories be? There sure can be sense in informing the public about community level decisions on construction and development based on town council decisions mined from public service databases. Inferring and precisely directing a message like “Town council planning bypass to relief your neighbourhood of cut-through traffic” can be relevant news to a small number of locals, but could well be too tedious and too low-impact for journalists to have to bother with. Automated services like that might thus be well placed actually.</p>

<p>The big problem why progress in developing such a service will be slow is that the closer you get to the community and the individual, the more heterogeneous and specific news needs and interests become. Inferring automatically that a plan for a motorway bypass will affect people in some area is one, but deciding what this <em>means</em> to people is a whole different ball game. One that machine learning is still terribly bad at. What open data sets are you going to use to have some sense of how to frame your news story, so important if it is to be tailored to local needs? Are we turning to what is available on the Web, for instance, produced by the community itself? I sincerely doubt that will result in actually “fact-based insights into local communities”, which is the “increasing demand” AP says to target. These are challenges of automated inference that are not easily solved, resulting in the slower bit: the output and impact of projects like these are usually far more modest—yet still usable—than press releases tend to suggest. A very nice prototype will be derived. Probably, maybe.</p>

<p>Another part of the slower bit is that it is easy enough to generate high level, mainstream interest, stock exchange news items from well formalised statistics, but that it is a lot harder to generate daily real life human interest stories. “First quarter income was reported at 15 million USD” is a sentence easily enough generated from well groomed statistics and databases. Generating “Bob’s farm house store will be closed temporarily in October” is of more immediate interest to some local community, but it involves complexities of automated inference from real world information far beyond current capabilities. Although generating language is getting <a href="https://chunml.github.io/ChunML.github.io/project/Creating-Text-Generator-Using-Recurrent-Neural-Network/">intriguingly easy</a>—as easy as predicting best sellers almost—it is still a far cry from the heterogeneous specificity that is relevant if you get to local or individual level. There is a meaningful and relevant difference between “Bob’s farm shop will be closed temporarily in October” and “Bob’s farm shop will be closing in October”. The first is an ordinary well formed announcement, the second is a potential source of hazardous assumptions about Bob and his commercial and physical well being. Current deep learning algorithms are rather insensitive to such subtleties and one has little control over whether one or the either will be generated. But such subtleties are what starts to matter if you get down to the less formulaic, less high level pattern based, nature of language in local community real life.</p>

<p>So in all it will be a while, and I would rather be interested in the results of the project than in shout outs about the potential demise of journalism writ large. At the same time I do not want to downplay the hazardous situation that these technologies may eventually put journalists in. This involves the even harder question of how we wield our digital technology ethically. Yes, it is possible to predict best sellers. Even the simplest possible application of deep learning yields 80% success. As a publisher you would be a rather poor businessman if you would not at least scout out the possible edge this might give you in a publishing industry. But that does not mean you need to do away with your editors. In fact that would be the intellectually poorest and most unethical choice. You can chose to do so, and you can even keep making a profit that way. The downside is however, that your predictive algorithm will force creative writing into a unpalatable sludge of same-plot-same-style novels. The more clever and ethical option is to use such an algorithm as a support tool to avert the worst potential non sellers. Each averted non seller—even if we get one in five wrong—is money saved to poor into more promising, more interesting projects. Responsibly deployed this algorithm enhances the ability of a publisher to support new and truly interesting work. An ethical businessman—I am aware this files as a modern oxymoron though—would turn losses prevented into an investment in new interesting literature that deviates form the by now plain old literary thriller genre. In machine learning that is something we are far worse at than humans still: predicting what outlier might actually not be a non seller, but an example of the new brilliant different style and narrative that will hit it big.</p>

<p>The rationale for how we choose to wield our technologies comes from critical thinking and reflection—or the lack thereof of course. It is the more deplorable therefore that PA’s press release resulted in a quite predictable flood of reactions in the “journalists’ jobs are under threat” genre. Because it means that journalists did mostly not do their jobs of critical news mongering. If they had, they had not talked so much about their jobs. For as argued, it will be a while until these are really under threat, if at all. Rather they could have talked about the highly arguable motives of an industry giant like Google being involved in developing algorithms that select and write news items. Do we really believe that these algorithms will be impartial? Will we yet again believe the stories that data science and natural language parsing algorithms are neutral because they are based on mathematics and logic? How likely is it that Google, more specifically its board, will be ethical about deploying these algorithms? That is the type of questions the press should have set out to answer. Instead of devoting time to being really investigative and critical, they chose to repeat the press release.</p>

<p>So don’t be surprised if you sometime soon read a local announcement: “Bob’s store will be closed temporarily in October, but you can find fine online stores on Google”.</p>

<p>–JZ_20170710_2017</p>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[Geert Lovink who does wonderful work at the Institute of Network Cultures yesterday tweeted a cry of horror on finding out via TechCrunch that Google is funding the development of software that writes local news stories. Media the world over have parroted the same news which seems largely based on a press release from the UK Press Association. The parroting in itself is a indicator of the dire situation in journalism where uncritically posting press releases has become a stand in for actual in depth and well researched coverage. Those who at least attempted a stab at a perspective mostly seem to have stuck to the hackneyed criticism that Google funds the development of robot journalists that will put human journalists out of a job: “Journalists, look out: Google is funding the rise of the AI news machine”.]]></summary></entry><entry><title type="html">Singularity</title><link href="http://localhost:4000/singularity/" rel="alternate" type="text/html" title="Singularity" /><published>2016-06-27T14:00:20+02:00</published><updated>2016-06-27T14:06:53+02:00</updated><id>http://localhost:4000/singularity</id><content type="html" xml:base="http://localhost:4000/singularity/"><![CDATA[<p>Willard McCarty on <a href="http://lists.digitalhumanities.org/pipermail/humanist/2016-June/013930.html">Humanist</a> pointed me to a, quite silly, <a href="http://www.economist.com/news/leaders/21701119-what-history-tells-us-about-future-artificial-intelligenceand-how-society-should">article</a> in the Economist entitled “March of the Machines”. It can almost be called a genre piece. The author downplays very much the possible negative effects of artificial intelligence and then argues that society should find an ‘intelligent response’ to AI—as opposed, I assume, to uninformed dystopian stories.</p>

<p>But I do hope the intelligent response society will seek to AI will be less intellectually lazy than the author of said contribution. I think to be honest that someone needed to crank out a 1000 words piece quickly, and reverted to sad stopgap rhetorics.</p>

<p>In this type of article there’s invariably a variation on this sentence: “Each time, in fact, technology ultimately created more jobs than it destroyed”. As if—not denying here any of a job’s power to be meaningful and fulfilling for many people—a job is the single quality of existence.</p>

<p>Worse is that such multi purpose filler arguments ignore unintended side effects of technological development. Mass production was brought on by mechanisation. We know that it also brought mass destruction. It is always sensible to consider both the possible dystopian and utopian scenarios. No matter what <a href="http://www.andrewng.org/">Andrew Ng</a> (quoted in the article) obviously should say as an AI researcher, it is actually very sensible to consider overpopulation of Mars before you colonise it. Before conditions are improved for human live there—at whatever expense—even a few persons will effectively establish such an overpopulation. Ng’s argument is a non sequitur anyway. If the premise of the article is correct we are not decades away from ubiquitous application of AI. Quite the opposite, the conditions on Earth for AI have been very favourable for more than a decade already. We hardly can wait to try out all our new toys.</p>

<p>No doubt AI will bring some good, and also no doubt it will bring a lot of awful bad. This is not inherent in the technology, but the in the people that wield it. Thus it is useful to keep critically examining all applications of all technologies while we develop them, instead of downplaying without evidence its unintended side effects.</p>

<p>If we do not, we may create our own foolish utopian illusions. For instance when we start using arguments such as “AI may itself help, by personalising computer-based learning and by identifying workers’ skills gaps and opportunities for retraining.” Which effectively means asking the machines what the machines think the non-machines should do. Well, if you ask a machine, chances are you’ll get a machinery answer and eventually a machinery society. Which might be fine for all I know, but I’d like that to be a very well informed choice.</p>

<p>I am not a believer of <a href="https://en.wikipedia.org/wiki/Technological_singularity">The Singularity</a>. Chances that machines and AI will aggressively push out human kind are in all likelihood gross exaggerations. But a realistic possibility is the covert permeation of human society by AI. We change society by our use of technology and the technology changes us too. This has been and will always be the case, and it is far from some moral or ethical wrong. But of these changes we should be conscious and informed, so that we hold the choice and not the machine. If a dialogue between man and (semi-)intelligent machine would be started as naive as the author of the Economist piece suggests, then human kind might indeed be very naively set to become machine like.</p>

<p>Machines and AI are, certainly until now, extensions and models of human behaviour. They are models and simulations of such behaviour, they are never humans. This can improve human existence manyfold. But having the heater on is something quite different than asking a model of yourself: “What gives my life meaning? How should I come to a fulfilling existence?” Asking that of a machine, even a very intelligent one, is still asking a <em>machine </em>what it is to be <em>human</em>. It is not at all excluded that a machine will not ever find a reasonable or valuable answer to that. But I would certainly wait beyond the first few iterations of this technology before possibly buying into any of the answers we might get.</p>

<p>It is deceptively easy to be unaware of such influences. In 1995 most people found cell phones marginally useful and far too expensive. A mere 20 years later almost no one wants to depart from his or her smartphone. This has changed how we communicate, when we communicate, how we live, who we are. AI will have similar consequences. Those might be good, those might be bad. They shouldn’t be however covert.</p>

<p>Thus I am not saying at all that a machine should never enter a dialogue with humans on human existence. But when we enter that dialogue we change the character of the interaction we have had with technologies since we can remember considerably. Humans have always defined technology, and our use of it has in part defined us. By changing technology we change ourselves. This acts out on the individual level—I am a different person now due to using programming languages than I was when I did not—and on the scale of society where we are part of socio-technical ecosystems comprising both technologies, communities, and individuals.</p>

<p>But these interactions have always been a monologue on the intellectual level. As soon as this becomes a dialogue because the technology literally can now speak to us, we need to be aware that it is not a human speaking to us, but a model of a human.</p>

<p>I for one would be excited to learn what that means, what riches is may bring. But I would always enter such a conversation well aware that I am talking not to another, but to a machine, and I would weigh that fact into the value and evaluation of the conversation. To assume that AI will answer questions on what course of action would lead me to improving my skills and my being, may be too heavily a buy in into the abilities of AI models to understand human life.</p>

<p>Sure AI can help. Even <em>more</em> so if we are aware of the fact that its helpful qualities are by definition limited to the realm of what the machine can understand.</p>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[Willard McCarty on Humanist pointed me to a, quite silly, article in the Economist entitled “March of the Machines”. It can almost be called a genre piece. The author downplays very much the possible negative effects of artificial intelligence and then argues that society should find an ‘intelligent response’ to AI—as opposed, I assume, to uninformed dystopian stories.]]></summary></entry><entry><title type="html">Methodological safety pin</title><link href="http://localhost:4000/methodological-safety-pin/" rel="alternate" type="text/html" title="Methodological safety pin" /><published>2016-06-20T11:30:18+02:00</published><updated>2016-06-20T11:30:18+02:00</updated><id>http://localhost:4000/methodological-safety-pin</id><content type="html" xml:base="http://localhost:4000/methodological-safety-pin/"><![CDATA[<p>There is a trope in digital humanities related articles that I find particularly awkward. Just now I stumbled across another example, and maybe it is a good thing to muze about it a short bit. Whence the example comes I don’t think is important as I am interested in the trope in general and not in this particular instance per sé. Besides, I like the authors and have nothing against their research, but before you know it flames are flying everywhere. So in the interest of all I file this one for prosperity anonymized.</p>

<p>This is the quote in question: “The first step towards the development of an open-source mining technology that can be used by historians without specific computer skills is to obtain a hands-on experience with research groups that use currently available open-source mining tools.”</p>

<p>Readers of digital humanities related essays, articles, reviews etc. will have found ample variations on this theme in the literature. From where I am sitting such statements rig up a dangerous strawman or facade. There are a number of hidden (and often not so hidden at all) assumptions that are glossed over with such statements.</p>

<p>First of all there is the assumption that it is obvious that as a scholar without specific computer skills you still should be able to use computer technology. This is a nice democratic principle I guess, but is it a wise one too?</p>

<p>Second, there’s the suggestion that all computer technology is homogeneous. There is no need to differentiate between levels and types of interfaces and technologies. It can all sweepingly be nicely represented as this amorphous mass of “open-source mining technology”. I know it is not entirely fair to pin this on the authors of such statements. Indeed the authors may be very well aware that they are generalizing a bit in service of the less experienced reader. However, the scholarly equivalent would be to say that the first step for a computer scientist that wants to understand history is to get a hands-on experience with historians. Even if that might be in general true, from scholarly arguing I expect more precision. You do not ‘understand history’. One understands tiny, very specific parts of it, maybe, when approached with very specific very narrowly formulated research questions, and meticulous methodology. I do not understand why the wide brush is suddenly allowed if the methodology turns digital.</p>

<p>Third, and this is the assumption that I find most problematic: there is the assumption (rather axiom maybe) that there shall be a middle man, a bridge builder, a guide, a mediator, or go-in-between that shall translate the expertise from the computer skilled persons involved towards the scholar. You hardly ever read it the other way round by the way, it is never the computer scientist in need of some scholarly wisdom. This in particular is a reflex and a trope I do not understand. When you need expertise you talk to the expert, and you try to acquire the expertise. But when it comes to computational expertise we (scholars) are suddenly in need of a mediator. Someone who goes in between and translates between expertises. In much literature—that in itself is part of this process of expertise exchange—this is now a sine qua non that does not get questioned at all: of course you do not talk to the experts directly, and of course you do not engage with the technology directly. When your car stalls, you don’t dive into the motor compartiment with your scholarly hands do you?!</p>

<p>Maybe not—though I at least try to determine even with my limited knowledge of car engines what might be the trouble. But I sure a hell talk to the expert directly. The mechanic is going to fix my car, I want to know what the trouble is and what he is going to do. Yes well, the scholar retorts, but quite frankly I do not talk so much on the car engine trouble to my mechanic at all! Fair enough, might not be your cup of tea. But the methodology of your research should be. Suppose you are diagnosed with cancer, do you want to talk only to the secretary of your doctor?</p>

<p>Besides, it <em>is</em> about the skills. A standard technique to disguise logical fallacies in reasoning is to substitute object phrases. I play this little game with these tropes too: “The first step towards the development of a hand grenade that can be used by historians without specific combat skills is to obtain a hands-on experience with soldiers that use currently available hand grenades.”</p>

<p>This doesn’t invalidate the general truthiness of the logic, but it does serve to lay bare its methodological fallacy: if you want to use that technology, better acquire some basic skills from the experts if you want to rely safely on the outcome of its use.</p>]]></content><author><name>Joris van Zundert</name><email>joris.van.zundert@gmail.com</email></author><summary type="html"><![CDATA[There is a trope in digital humanities related articles that I find particularly awkward. Just now I stumbled across another example, and maybe it is a good thing to muze about it a short bit. Whence the example comes I don’t think is important as I am interested in the trope in general and not in this particular instance per sé. Besides, I like the authors and have nothing against their research, but before you know it flames are flying everywhere. So in the interest of all I file this one for prosperity anonymized.]]></summary></entry></feed>