<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Fennix</title>
    <link>https://infosec.press/fennix/</link>
    <description></description>
    <pubDate>Wed, 16 Sep 2026 21:22:42 +0000</pubDate>
    <item>
      <title>The state of web ads is over a century in the making</title>
      <link>https://infosec.press/fennix/the-state-of-web-ads-is-over-a-century-in-the-making</link>
      <description>&lt;![CDATA[Preface: I was originally going to go on a rant but fell down a rabbit hole of looking at examples of older newspapers and instead this became more of an article/blog.&#xA;&#xA;Like many other people focused on their privacy, I run Pi-hole at home to block advertising domains, among other annoyances, and personally make extensive use of Privacy Badger and NoScript.  The Pi-hole alone has the effect whenever anyone of the household is out of the building and not connected to our home&#39;s wifi, they get the jarring experience of seeing a completely different version of the web, plastered with ads, most of which are animated and attention-grabbing. This is especially true in mobile apps.&#xA;&#xA;There&#39;s been a lot of lip service given to the way the &#34;old web&#34; used to appear versus how things are now, so I&#39;m not going to do more of that here.  I think what a lot of us &#34;web old-timers&#34; maybe don&#39;t realize is that how the web looks now is actually pretty common and has its origins in the way newspapers and magazines were laid out in the past.&#xA;&#xA;!--more--&#xA;&#xA;If you&#39;ve never had cause to go back and look at old newspaper archives, you might not have experienced this, so I&#39;m going to show you some examples.  You&#39;ll see the bones of modern web advertising buried in newspapers a century old, and then I&#39;ll explain what I think is critically different about how the web is these days.  Spoiler: it&#39;s not better.&#xA;&#xA;img src=&#34;https://pixel.infosec.exchange/storage/m/v2/521845472966472902/3079cad20-917577/jD3ub25VNdwu/VCaOQoE4ELv622rkIEuxJslWFMh03Aziy5xyZCcD.png&#34; alt=&#34;The front page of the New York Times,&#39; Tuesday, February 1, 1921 edition. It looks very different to modern day front pages of newspapers, being divided into 8 columns with a dozen or more stories visible. There are no ads present.&#34;&#xA;span style=&#34;font-size:0.70em;&#34;ref: https://archive.org/details/NYTimesfeb1151921/page/mode/2up/span&#xA;&#xA;You can see there&#39;s a few modern innovations in papers missing here -- no &#34;above the fold&#34; style of breaking up the layout.  Another trick newspapers did initially was to never put the ads on the front page -- they were selling you the news after all.&#xA;&#xA;Now onto page two:&#xA;&#xA;img src=&#34;https://pixel.infosec.exchange/storage/m/v2/521845472966472902/3079cad20-917577/sfrhh3f2E5Ju/Swk2bMuIE9BDAJSvlO9Onpe9nGy9yqUbI7O2ir5O.png&#34; alt=&#34;Page two of the New York Times,&#39; Tuesday, February 1, 1921 edition. From left to right the page is approximately divided 80/20 between real news and ads. There are between half a dozen and a dozen stories on the page.&#34;&#xA;span style=&#34;font-size:0.70em;&#34;ref: https://archive.org/details/NYTimesfeb1151921/page/n1/mode/2up/span&#xA;&#xA;This layout is common amongst the meatier news-focused sections of the paper. The first 5.5 columns are dedicated to news stories and then the rest is devoted to ads. Three are larger double-column spaced ads, while two are smaller and occupy the space in an ad. The rest of the Times&#39; early layouts in the news sections were like this, with sometimes more space dedicated to ads on the lighter topics. &#xA;&#xA;For example, here&#39;s the sports section:&#xA;img src=&#34;https://pixel.infosec.exchange/storage/m/v2/521845472966472902/3079cad20-917577/5vFMb37jFCBf/3F6oEt0kKZymjC8eBfZ9Nes6ZVVu7I6P7ERD3aoE.png&#34; alt=&#34;Page 12 of the New York Times,&#39; Tuesday, February 1, 1921 edition. This is the sports section. It&#39;s divided roughly 65/35 between ads and stories, and features early versions of a popular web advertising layout where the side columns are dedicated to advertising.&#34;&#xA;span style=&#34;font-size:0.70em;&#34;ref: https://archive.org/details/NYTimesfeb1151921/page/n11/mode/2up/span&#xA;&#xA;Here ads are placed in a very familiar format for the modern web; The ads effectively bookend either side of the center columns which house the articles themselves. &#xA;&#xA;However, it&#39;s worth noting this layout was not universal. Here&#39;s an example of the Victoria Daily Times, from Victoria, British Columbia:&#xA;img src=&#34;https://pixel.infosec.exchange/storage/m/v2/521845472966472902/3079cad20-917577/dsHXtRWat164/1LdIQiQIv3dUvZYHZ9JstpjOWuJuNvEO853QObjb.png&#34; alt=&#34;Page 2 of the Victoria Daily Times&#39; Friday July 22, 1921 edition. This uses a very different layout than the New York Times. Here ads are sometimes placed in the center columns breaking up the stories. There does not appear to be any standardized ad sizes either, beyond snapping to columns for width.&#34;&#xA;span style=&#34;font-size:0.70em;&#34;ref: https://archive.org/details/victoriadailytimes19210722/page/n1/mode/2up/span&#xA;&#xA;I can only imagine its print runs were much smaller than the New York Times. The paper lives on to this day as the Victoria Times Colonist, having merged with another local paper in the 1980s.  Attempting to read this layout now, I understand why the format the Times is using won out over other layouts. The ads being so close to the article is visually distracting.&#xA;&#xA;Now let&#39;s compare those older examples to modern web news media.&#xA;Let&#39;s start with a relatively tame example: Yahoo! News:&#xA;&#xA;img src=&#34;https://pixel.infosec.exchange/storage/m/v2/521845472966472902/3079cad20-917577/cDqhE4WRX6Bb/vNwXgeSYBgFFd49j3KHS6Y563aQ1hjDhEabCGy6y.png&#34; alt=&#34;Yahoo! News article titled &#39;Sister André — the world&#39;s oldest person — has died at 118. She drank a glass of wine every day and credited her long life to working until she was 108.&#39; published Wednesday, January 18, 2023. The article&#39;s author is listed as Rebecca Cohen. Ads are visible largely down the right column, mimicking the layout of the earlier 1920s New York Times. A similar ratio of space is devoted to the side ad bar as well, roughly 20%. Below the title text is a photo of Sister André, with her hands clasped in a prayer gesture, taken April 27, 2022.&#34;&#xA;span style=&#34;font-size:0.70em;&#34;ref: https://us.yahoo.com/news/sister-andr-worlds-oldest-person-183029182.html/span&#xA;&#xA;Here you can see the remnants of the earlier newspaper design. The page is divided into roughly fifths, and a fifth is allocated to the side ads.&#xA;All in all this doesn&#39;t look too unreasonable, but let&#39;s now look to what modern newspapers&#39; sites look like. Here&#39;s the front page of the New York times:&#xA;&#xA;img src=&#34;https://pixel.infosec.exchange/storage/m/v2/521845472966472902/3079cad20-917577/e3T2AzJYn3t9/xEBKcVBi5V5nQVqqXefCLeJ9Hb5ZMK4ErCw1yXue.png&#34; alt=&#34;Front page of the New York Times&#39; website, January 19, 2023 (20th in some locales). A large banner ad which failed to load occupies the top two fifths to one half of the visible page space, with stories below. Stories appear to have one primary column, occupying three quarters of the width, with the last quarter being devoted to other smaller articles.&#34;&#xA;span style=&#34;font-size:0.70em;&#34;ref: https://nytimes.com//span&#xA;&#xA;Because of the nature of the web and the drive to obtain impressions versus what works in print there&#39;s a huge functional difference: Each article is given its own webpage, so what is displayed on the main page landing page actually looks more like older newspapers, where a single viewing space -- in that time period, paper, in ours, screen real estate -- is subdivided into several articles. I don&#39;t have any insider knowledge or analytics but I believe that in today&#39;s social media dominated world, most users do not visit the front pages of newspaper sites, yet the philosophy persists of the relatively &#34;clean&#34; first page. &#xA;&#xA;Now let&#39;s look at what happens when we load an article:&#xA;&#xA;img src=&#34;https://pixel.infosec.exchange/storage/m/v2/521845472966472902/3079cad20-917577/axmxcMAeLEgk/ePwIYeCD0WmN7iZboE5xWsY51nu5TtSi8OPcMPtz.png&#34; alt=&#34;Screen shot of the New York Times&#39; article titled &#39;Supreme Court Says it Hasn&#39;t Found Who Leaked Opinion Overturning Roe&#39;, dated January 19, 2023. There is a large banner ad consuming approximately two fifths of the viewing space for a subscription service to the Times. The bottom half of the page is a popup message asking the user to create a free account or log in to continue reading articles. Only the very tops of the upper-cased letters in the article title are visible at all.&#34;&#xA;span style=&#34;font-size:0.70em;&#34;ref: https://www.nytimes.com/2023/01/19/us/politics/supreme-court-leak-roe.html/span&#xA;&#xA;Here we can see probably the worst feature of modern advertising: the pop up modal dialog requesting subscription or registration.  This is commonplace among all newspapers websites at this point in time and that&#39;s not news to any of you.  On what should be an article page the title is not even visible!  There are two separate subscribe buttons visible, plus our lovely &#34;create an account&#34; modal dialog.&#xA;&#xA;Not that I&#39;m unsympathetic, the trials of various news organizations are well documented so I don&#39;t need to go into them here.  What I would like to highlight is simply that some philosophies and design elements in use a hundred years ago persist.  For example, we still have the behavior of keeping the first point of arrival largely ad free.  Not completely of course, because the tombstone of the 2015+ web will be engraved with &#34;Subscribe, Click that like button, and share it with your friends&#34;, but the front page is relatively ad-free compared to the hilarious experience of trying to view an article.&#xA;&#xA;On that note though, the viewing an article experience is very reminiscent of the Victoria Daily Times&#39; layout. Maybe they were right all along.&#xA;&#xA;The problem this creates is that whenever I visit friends or family who aren&#39;t tech-savvy, I realize just how bombarded they get with advertising.&#xA;&#xA;It also really drives home Google&#39;s impetus for working on DNS-over-HTTPS and Manifest V3: It will help them take back control over ad visibility in the era of every user using an ad blocker in their browser and things like Pi-hole becoming cheaper and simpler for people to run at home.]]&gt;</description>
      <content:encoded><![CDATA[<p>Preface: I was originally going to go on a rant but fell down a rabbit hole of looking at examples of older newspapers and instead this became more of an article/blog.</p>

<p>Like many other people focused on their privacy, I run Pi-hole at home to block advertising domains, among other annoyances, and personally make extensive use of Privacy Badger and NoScript.  The Pi-hole alone has the effect whenever anyone of the household is out of the building and not connected to our home&#39;s wifi, they get the jarring experience of seeing a completely different version of the web, plastered with ads, most of which are animated and attention-grabbing. This is especially true in mobile apps.</p>

<p>There&#39;s been a lot of lip service given to the way the “old web” used to appear versus how things are now, so I&#39;m not going to do more of that here.  I think what a lot of us “web old-timers” maybe don&#39;t realize is that how the web looks now is actually pretty common and has its origins in the way newspapers and magazines were laid out in the past.</p>



<p>If you&#39;ve never had cause to go back and look at old newspaper archives, you might not have experienced this, so I&#39;m going to show you some examples.  You&#39;ll see the bones of modern web advertising buried in newspapers a century old, and then I&#39;ll explain what I think is critically different about how the web is these days.  Spoiler: it&#39;s not better.</p>

<p><img src="https://pixel.infosec.exchange/storage/m/_v2/521845472966472902/3079cad20-917577/jD3ub25VNdwu/VCaOQoE4ELv622rkIEuxJslWFMh03Aziy5xyZCcD.png" alt="The front page of the New York Times,&#39; Tuesday, February 1, 1921 edition. It looks very different to modern day front pages of newspapers, being divided into 8 columns with a dozen or more stories visible. There are no ads present.">
<span style="font-size:0.70em;">ref: <a href="https://archive.org/details/NYTimes_feb1_15_1921/page/mode/2up" rel="nofollow">https://archive.org/details/NYTimes_feb1_15_1921/page/mode/2up</a></span></p>

<p>You can see there&#39;s a few modern innovations in papers missing here — no “above the fold” style of breaking up the layout.  Another trick newspapers did initially was to never put the ads on the front page — they were selling you the news after all.</p>

<p>Now onto page two:</p>

<p><img src="https://pixel.infosec.exchange/storage/m/_v2/521845472966472902/3079cad20-917577/sfrhh3f2E5Ju/Swk2bMuIE9BDAJSvlO9Onpe9nGy9yqUbI7O2ir5O.png" alt="Page two of the New York Times,&#39; Tuesday, February 1, 1921 edition. From left to right the page is approximately divided 80/20 between real news and ads. There are between half a dozen and a dozen stories on the page.">
<span style="font-size:0.70em;">ref: <a href="https://archive.org/details/NYTimes_feb1_15_1921/page/n1/mode/2up" rel="nofollow">https://archive.org/details/NYTimes_feb1_15_1921/page/n1/mode/2up</a></span></p>

<p>This layout is common amongst the meatier news-focused sections of the paper. The first 5.5 columns are dedicated to news stories and then the rest is devoted to ads. Three are larger double-column spaced ads, while two are smaller and occupy the space in an ad. The rest of the Times&#39; early layouts in the news sections were like this, with sometimes more space dedicated to ads on the lighter topics.</p>

<p>For example, here&#39;s the sports section:
<img src="https://pixel.infosec.exchange/storage/m/_v2/521845472966472902/3079cad20-917577/5vFMb37jFCBf/3F6oEt0kKZymjC8eBfZ9Nes6ZVVu7I6P7ERD3aoE.png" alt="Page 12 of the New York Times,&#39; Tuesday, February 1, 1921 edition. This is the sports section. It&#39;s divided roughly 65/35 between ads and stories, and features early versions of a popular web advertising layout where the side columns are dedicated to advertising.">
<span style="font-size:0.70em;">ref: <a href="https://archive.org/details/NYTimes_feb1_15_1921/page/n11/mode/2up" rel="nofollow">https://archive.org/details/NYTimes_feb1_15_1921/page/n11/mode/2up</a></span></p>

<p>Here ads are placed in a very familiar format for the modern web; The ads effectively bookend either side of the center columns which house the articles themselves.</p>

<p>However, it&#39;s worth noting this layout was not universal. Here&#39;s an example of the Victoria Daily Times, from Victoria, British Columbia:
<img src="https://pixel.infosec.exchange/storage/m/_v2/521845472966472902/3079cad20-917577/dsHXtRWat164/1LdIQiQIv3dUvZYHZ9JstpjOWuJuNvEO853QObjb.png" alt="Page 2 of the Victoria Daily Times&#39; Friday July 22, 1921 edition. This uses a very different layout than the New York Times. Here ads are sometimes placed in the center columns breaking up the stories. There does not appear to be any standardized ad sizes either, beyond snapping to columns for width.">
<span style="font-size:0.70em;">ref: <a href="https://archive.org/details/victoriadailytimes19210722/page/n1/mode/2up" rel="nofollow">https://archive.org/details/victoriadailytimes19210722/page/n1/mode/2up</a></span></p>

<p>I can only imagine its print runs were much smaller than the New York Times. The paper lives on to this day as the <a href="https://www.timescolonist.com/" rel="nofollow">Victoria Times Colonist</a>, having merged with another local paper in the 1980s.  Attempting to read this layout now, I understand why the format the Times is using won out over other layouts. The ads being so close to the article is visually distracting.</p>

<p>Now let&#39;s compare those older examples to modern web news media.
Let&#39;s start with a relatively tame example: Yahoo! News:</p>

<p><img src="https://pixel.infosec.exchange/storage/m/_v2/521845472966472902/3079cad20-917577/cDqhE4WRX6Bb/vNwXgeSYBgFFd49j3KHS6Y563aQ1hjDhEabCGy6y.png" alt="Yahoo! News article titled &#39;Sister André — the world&#39;s oldest person — has died at 118. She drank a glass of wine every day and credited her long life to working until she was 108.&#39; published Wednesday, January 18, 2023. The article&#39;s author is listed as Rebecca Cohen. Ads are visible largely down the right column, mimicking the layout of the earlier 1920s New York Times. A similar ratio of space is devoted to the side ad bar as well, roughly 20%. Below the title text is a photo of Sister André, with her hands clasped in a prayer gesture, taken April 27, 2022.">
<span style="font-size:0.70em;">ref: <a href="https://us.yahoo.com/news/sister-andr-worlds-oldest-person-183029182.html" rel="nofollow">https://us.yahoo.com/news/sister-andr-worlds-oldest-person-183029182.html</a></span></p>

<p>Here you can see the remnants of the earlier newspaper design. The page is divided into roughly fifths, and a fifth is allocated to the side ads.
All in all this doesn&#39;t look too unreasonable, but let&#39;s now look to what modern newspapers&#39; sites look like. Here&#39;s the front page of the New York times:</p>

<p><img src="https://pixel.infosec.exchange/storage/m/_v2/521845472966472902/3079cad20-917577/e3T2AzJYn3t9/xEBKcVBi5V5nQVqqXefCLeJ9Hb5ZMK4ErCw1yXue.png" alt="Front page of the New York Times&#39; website, January 19, 2023 (20th in some locales). A large banner ad which failed to load occupies the top two fifths to one half of the visible page space, with stories below. Stories appear to have one primary column, occupying three quarters of the width, with the last quarter being devoted to other smaller articles.">
<span style="font-size:0.70em;">ref: <a href="https://nytimes.com/" rel="nofollow">https://nytimes.com/</a></span></p>

<p>Because of the nature of the web and the drive to obtain impressions versus what works in print there&#39;s a huge functional difference: Each article is given its own webpage, so what is displayed on the main page landing page actually looks more like older newspapers, where a single viewing space — in that time period, paper, in ours, screen real estate — is subdivided into several articles. I don&#39;t have any insider knowledge or analytics but I believe that in today&#39;s social media dominated world, most users do not visit the front pages of newspaper sites, yet the philosophy persists of the relatively “clean” first page.</p>

<p>Now let&#39;s look at what happens when we load an article:</p>

<p><img src="https://pixel.infosec.exchange/storage/m/_v2/521845472966472902/3079cad20-917577/axmxcMAeLEgk/ePwIYeCD0WmN7iZboE5xWsY51nu5TtSi8OPcMPtz.png" alt="Screen shot of the New York Times&#39; article titled &#39;Supreme Court Says it Hasn&#39;t Found Who Leaked Opinion Overturning Roe&#39;, dated January 19, 2023. There is a large banner ad consuming approximately two fifths of the viewing space for a subscription service to the Times. The bottom half of the page is a popup message asking the user to create a free account or log in to continue reading articles. Only the very tops of the upper-cased letters in the article title are visible at all.">
<span style="font-size:0.70em;">ref: <a href="https://www.nytimes.com/2023/01/19/us/politics/supreme-court-leak-roe.html" rel="nofollow">https://www.nytimes.com/2023/01/19/us/politics/supreme-court-leak-roe.html</a></span></p>

<p>Here we can see probably the worst feature of modern advertising: the pop up modal dialog requesting subscription or registration.  This is commonplace among all newspapers websites at this point in time and that&#39;s not news to any of you.  On what should be an article page the title is not even visible!  There are two separate subscribe buttons visible, plus our lovely “create an account” modal dialog.</p>

<p>Not that I&#39;m unsympathetic, the trials of various news organizations are well documented so I don&#39;t need to go into them here.  What I would like to highlight is simply that some philosophies and design elements in use a hundred years ago persist.  For example, we still have the behavior of keeping the first point of arrival largely ad free.  Not completely of course, because the tombstone of the 2015+ web will be engraved with “Subscribe, Click that like button, and share it with your friends”, but the front page is relatively ad-free compared to the hilarious experience of trying to view an article.</p>

<p>On that note though, the viewing an article experience is very reminiscent of the Victoria Daily Times&#39; layout. Maybe they were right all along.</p>

<p>The problem this creates is that whenever I visit friends or family who aren&#39;t tech-savvy, I realize just how bombarded they get with advertising.</p>

<p>It also really drives home Google&#39;s impetus for working on DNS-over-HTTPS and Manifest V3: It will help them take back control over ad visibility in the era of every user using an ad blocker in their browser and things like Pi-hole becoming cheaper and simpler for people to run at home.</p>
]]></content:encoded>
      <guid>https://infosec.press/fennix/the-state-of-web-ads-is-over-a-century-in-the-making</guid>
      <pubDate>Fri, 20 Jan 2023 00:18:25 +0000</pubDate>
    </item>
    <item>
      <title>Rantuary 15, 2023</title>
      <link>https://infosec.press/fennix/rantuary-15-2023</link>
      <description>&lt;![CDATA[Hello everybody.  My nick is Fennix, I&#39;m an app breaker by day and night. I might make this a daily thing I might make this every few days I am not sure yet.&#xA;&#xA;For today&#39;s rant I want to talk about libraries, their developers, and when not applying the Unix philosophy goes terribly wrong.&#xA;&#xA;I&#39;m going to talk about Log4J but I&#39;m also going to talk about things like XXE and in general design choices that lead to headaches. &#xA;&#xA;!--more--&#xA;&#xA;When you&#39;re designing a library that is intended to be used to tackle some important but common function, it&#39;s incredibly important that you keep the library as task focused as possible especially the core library and its defaults. If you need to extend functionality, use a pluggable architecture and make those plugins opt-in.  The amount of headache that Log4J (the &#34;log4shell&#34; vulnerability really) caused the world is outsized to what everyone expected the library to do.&#xA;&#xA;It&#39;s important to understand that users&#39; expectations of what the library is doing are important. Log4J is not alone in this though.  The log4shell vulnerability is very reminiscent to me of XXE. It&#39;s a feature that was enabled as a default to do some additional parsing that most of its users didn&#39;t want or need and that they didn&#39;t necessarily have visibility to.&#xA;&#xA;Along those lines, if you&#39;re not familiar with XXE, AKA XML External Entity parsing attacks, the basics of the attack are this:&#xA;Attacker submits XML to server&#xA;Server parses XML&#xA;Server does a bunch of stupid shit like opening remote connections and sending files&#xA;Attacker laughs, possibly even a good cackle&#xA;&#xA;When XML as a document standard was being ratified importance was placed on this idea of being able to validate the document against an arbitrary schema in order to make it flexible.  It was important that schema specifications not just be able to be loaded from local files but could be loaded from central locations using a variety of different protocols. Examples of these are Gopher, FTP, or later HTTP.  XML is very old.&#xA;&#xA;Secondly, in XML there is this concept of entities -- a shorthand within the document so that you can refer to some special character or a predefined standard blurb. You have likely seen these; the &amp;copy; that you would use to insert a copyright symbol in an older HTML doc is an entity (HTML having its roots in XML).  When you combine these two things what it meant is that you could have remotely loadable entities that would get parsed and loaded on the machine that was processing the document.&#xA;&#xA;Now because you might have some rather large entity, perhaps some boilerplate legalese that needs to be attached to each document, you might want to load that out of a local text file. You might make &amp;legalese; into an entity that reads its data from /usr/lib/standard_disclaimer.txt.&#xA;&#xA;This idea of document processor went from simple to unfocused, and because of these features you can probably see how with XXE you could often steal contents of files, reveal remote server locations, SSRF, cause a denial of service, etc., purely because this specification became overly complicated. &#xA;&#xA;It was then made worse by the fact that as the web was evolving, nobody had a better answer than XML for a long time to do online document exchange. Since it was already a standard in business, it meant that it had the inertia and so there was no reason to change this.  Ultimately you end up with major websites being vulnerable to all manner of XXE attacks purely because some support for some long forgotten feature was thrown in there. Even today this happens.&#xA;&#xA;Enter the developer using it: It&#39;s not clear that this needs to be turned off, I just wanted to parse an XML document!  They don&#39;t make any mention of this sort of thing anywhere in the documentation, so why would I think it&#39;s by nature unsafe?!&#xA;&#xA;This is horse shit.  We should expect better from our library authors.  We should expect better from commonly used components.  Importantly all the billion dollar corporations that make much of their billions leveraging this kind of software need to pony up, fund some pentests for these things, fund developer education, dedicate some resources to it.]]&gt;</description>
      <content:encoded><![CDATA[<p>Hello everybody.  My nick is Fennix, I&#39;m an app breaker by day and night. I might make this a daily thing I might make this every few days I am not sure yet.</p>

<p>For today&#39;s rant I want to talk about libraries, their developers, and when not applying the Unix philosophy goes terribly wrong.</p>

<p>I&#39;m going to talk about Log4J but I&#39;m also going to talk about things like XXE and in general design choices that lead to headaches.</p>



<p>When you&#39;re designing a library that is intended to be used to tackle some important but common function, it&#39;s incredibly important that you keep the library as task focused as possible especially the core library and its defaults. If you need to extend functionality, use a pluggable architecture and make those plugins opt-in.  The amount of headache that Log4J (the “log4shell” vulnerability really) caused the world is outsized to what everyone expected the library to do.</p>

<p>It&#39;s important to understand that users&#39; expectations of what the library is doing are important. Log4J is not alone in this though.  The log4shell vulnerability is very reminiscent to me of XXE. It&#39;s a feature that was enabled as a default to do some additional parsing that most of its users didn&#39;t want or need and that they didn&#39;t necessarily have visibility to.</p>

<p>Along those lines, if you&#39;re not familiar with XXE, AKA XML External Entity parsing attacks, the basics of the attack are this:
– Attacker submits XML to server
– Server parses XML
– Server does a bunch of stupid shit like opening remote connections and sending files
– Attacker laughs, possibly even a good cackle</p>

<p>When XML as a document standard was being ratified importance was placed on this idea of being able to validate the document against an arbitrary schema in order to make it flexible.  It was important that schema specifications not just be able to be loaded from local files but could be loaded from central locations using a variety of different protocols. Examples of these are Gopher, FTP, or later HTTP.  XML is very old.</p>

<p>Secondly, in XML there is this concept of entities — a shorthand within the document so that you can refer to some special character or a predefined standard blurb. You have likely seen these; the <code>&amp;copy;</code> that you would use to insert a copyright symbol in an older HTML doc is an entity (HTML having its roots in XML).  When you combine these two things what it meant is that you could have remotely loadable entities that would get parsed and loaded on the machine that was processing the document.</p>

<p>Now because you might have some rather large entity, perhaps some boilerplate legalese that needs to be attached to each document, you might want to load that out of a local text file. You might make <code>&amp;legalese;</code> into an entity that reads its data from <code>/usr/lib/standard_disclaimer.txt</code>.</p>

<p>This idea of document processor went from simple to unfocused, and because of these features you can probably see how with XXE you could often steal contents of files, reveal remote server locations, SSRF, cause a denial of service, etc., purely because this specification became overly complicated.</p>

<p>It was then made worse by the fact that as the web was evolving, nobody had a better answer than XML for a long time to do online document exchange. Since it was already a standard in business, it meant that it had the inertia and so there was no reason to change this.  Ultimately you end up with major websites being vulnerable to all manner of XXE attacks purely because some support for some long forgotten feature was thrown in there. Even today this happens.</p>

<p>Enter the developer using it: It&#39;s not clear that this needs to be turned off, I just wanted to parse an XML document!  They don&#39;t make any mention of this sort of thing anywhere in the documentation, so why would I think it&#39;s by nature unsafe?!</p>

<p>This is horse shit.  We should expect better from our library authors.  We should expect better from commonly used components.  Importantly all the billion dollar corporations that make much of their billions leveraging this kind of software need to pony up, fund some pentests for these things, fund developer education, dedicate some resources to it.</p>
]]></content:encoded>
      <guid>https://infosec.press/fennix/rantuary-15-2023</guid>
      <pubDate>Mon, 16 Jan 2023 03:41:34 +0000</pubDate>
    </item>
  </channel>
</rss>